איתי ספרן

אקדמי בכיר

Depth Separations in Neural Networks

Separating the Dimension from the Accuracy

Itay Safran, Daniel Reichman, Paul Valiant

We prove an exponential separation between depth 2 and depth 3 neural networks, when approximating a O(1)-Lipschitz target function to constant accuracy, with respect to a distribution with support in the unit ball, under the mild assumption that the weights of the depth 2 network are exponentially bounded. This resolves an open problem posed in Safran et al. (2019), and proves that the curse of dimensionality manifests itself in depth 2 approximation, even in cases where the target function can be represented efficiently using a depth 3 network. Previously, lower bounds that were used to separate depth 2 from depth 3 networks required that at least one of the Lipschitz constant, target accuracy or (some measure of) the size of the domain of approximation scale polynomially with the input dimension, whereas in our result these parameters are fixed to be constants independent of the input dimension: our parameters are simultaneously optimal. Our lower bound holds for a wide variety of activation functions, and is based on a novel application of a worst-to average-case random self-reducibility argument, allowing us to leverage depth 2 threshold circuits lower bounds in a new domain.

שפת פרסום אנגלית
כתב עת Proceedings of Machine Learning Research
כרך 291
סטטוס פרסום פורסם - 01.01.2025

Keywords

Depth separations
Function approximation
Lower bounds
Neural networks

ASJC Scopus subject areas

Software
Control and Systems Engineering
Statistics and Probability
Artificial Intelligence
קבצים וקישורים אחרים
Link to publication in Scopus