Itay Safran

Senior Academic

Provable Privacy Attacks on Trained Shallow Neural Networks

Guy Smorodinsky, Gal Vardi, Itay Safran

We study what provable privacy attacks can be shown for trained 2-layer ReLU neural networks, focusing on two types of attacks: membership inference and data reconstruction. We prove that theoretical results on the implicit bias of 2-layer neural networks can be used to provably identify with high probability whether a given point was used in the training set in a high-dimensional, nearly orthogonal setting, and can also be used to construct a finite set of which at least a constant fraction are training points in a univariate setting. To the best of our knowledge, our work is the first to show provable vulnerabilities in this implicit-bias-driven setting.

Publication language English
Journal Transactions on Machine Learning Research
Volume 2026-August
Publication status Published - 01.01.2026

ASJC Scopus subject areas

Computer Vision and Pattern Recognition
Artificial Intelligence
Other files and links
Link to publication in Scopus