Test error of interpolating two-layer networks in the feature-learning regime
In plain words
Networks with far more adjustable weights than training examples can fit noisy data exactly and still predict new data well, against the classical expectation of overfitting (memorizing noise). The task is to compute the error on new data for networks whose internal features adapt during training.
Precise statement
Two-layer network $f(x) = \operatorname{sum}_{j=1..m} a_j \sigma(w_j . x)$ trained by gradient descent from small initialization to zero training loss on n samples $y = f_{*}(x) + \mathrm{noise}$ of variance $\sigma_n^2$, $x \sim \mathrm{N}(0, I_d)$ with $I_d$ the d x d identity, in the limit $n, d \to \infty$ with $\psi = n/d$ fixed and $m >> n$. For targets $f_{*}(x) = g(U x)$ with $U$ an r x d matrix and r fixed, compute the limiting test mean-squared error on noisy labels $E(\psi, \sigma_n^2, g)$ at fixed $\psi$, and classify the interpolation by the $\psi \to \infty$ limit of $E$: benign if $E \to \sigma_n^2$, tempered if $E \to c\,\sigma_n^2$ with $1 < c < \infty$, catastrophic if $E$ diverges. The answer is an asymptotic formula, proved or confirmed numerically to high precision.
What would settle it
An exact asymptotic formula for the test error of gradient-trained interpolating two-layer networks in the feature-learning regime, matched by large-scale simulation.
Status in the literature
Unverified note
Exact answers exist for linear regression (Bartlett, Long, Lugosi and Tsigler 2020) and random-feature models (Mei and Montanari 2022), a benign, tempered, catastrophic taxonomy was proposed by Mallinar et al. 2022, and partial feature-learning results exist for classification and low-dimensional settings (Frei, Chatterji and Bartlett 2022; Kornowski, Yehudai and Shamir 2023; Joshi, Vardi and Srebro 2023); no asymptotic test-error formula for gradient-trained interpolating networks in the proportional regime exists (2026).