Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think the idea proposed is that if you add white noise on top of the adversarial perturbation, it will destroy that very particular perturbation.


It won't destroy the perturbation. There is a low dimensional manifold along which the cost function will decrease/increase. The adversarial perturbation lies on that manifold, but random white noise (which is a random high dimensional vector) will have close to zero length with high probability when projected onto the manifold and hence won't affect the cost function.


Ohhhh, I see. I'm not sure whether that would work. I would have to actually do the experiment. It may not work because this bizarro region of parameter space may be somewhat robust to perturbations, so in a sense you may have to travel out same way you travelled in. But, then again, maybe not.

Edit: I should add that these perturbations appear to be very robust to different architectures and datasets. So, the same adversarial perturbation will trick different NNs that were trained on different datasets. This suggests that it will probably be fairly robust to noise. But maybe not! I'm not aware of this experiment having been done.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: