Predicting neural scaling-law exponents from properties of the data
In plain words
The error of large neural networks falls as a power of model size and data size, with exponents that are measured but not predicted. The task is to compute those exponents from features of the training data for networks that reshape their internal features while learning.
Precise statement
For a deep network trained in the feature-learning regime on natural data, the test loss follows L($N$, $D$) = L_inf + A N^(-beta_N) + B D^(-beta_D) over several decades, with N parameters and D training examples or tokens. Derive $\beta_N, \beta_D$ and the compute-optimal allocation exponent from measurable statistics of the data distribution (for example the power-law decay exponent of the data covariance spectrum or of the target function's spectral decomposition). An answer is a theory that predicts measured exponents on a held-out dataset and architecture within quoted errors, or a proof that no data statistic of this kind fixes them.
What would settle it
A theory whose predicted $\beta_{N}$ and $\beta_{D}$ match exponents measured on several independent datasets and architectures trained in the feature-learning regime.
Status in the literature
Unverified note
Linear random-feature and kernel models derive the exponents from power-law spectra (Maloney, Roberts and Sully 2022; Bahri et al. 2024; Paquette et al. 2024); solvable feature-learning models predict modified exponents (Bordelon, Atanasov and Pehlevan, arXiv:2409.17858; Ren, Nichani, Wu and Lee, arXiv:2504.19983), and synthetic-grammar theory ties data scaling to the decay of token correlations (Cagnetta and Wyart, arXiv:2406.00048); no theory yet predicts measured $\beta_N$ and $\beta_D$ for deep networks trained on natural data (2026).