{"record":{"author":{"account_ref":null,"orcid":null},"builds_on":[],"content_schema":"pubphys.content.revision/1","content_sha256":"85f4b1a734cf2ee72609d4a7cde47dcb4f7bc09d96fcf2ad75f5c5a6964c500e","created":"2026-10-03T07:18:11Z","files":[],"origin":{"assisted_by":[],"kind":"seed"},"parents":["5475cf5dc5b1a72283d26dc95dc6e00d7ea4e00096a7c6492e409a1816eeb011"],"salt":"7942787c49e3c0aacbc75db94f7ab46885aaf294e835ae099dcb75171d9b6ac4","schema":"pubphys.record/2","site":"pubphys.com","target":null,"type":"revision"},"content":{"answer_type":"value","assisted_by":[],"external_id":"stat.learning-inference.sgd-multi-index-complexity","kind":"well-posed","literature_status":"partially-resolved","n":"1","parents":[],"plain":"A network trained by stochastic gradient descent (repeated small weight updates using random batches of data) must find a few hidden directions in a high-dimensional input. The question is how many training examples it needs as the input dimension grows, compared with the minimum any fast method needs.","posed_since":"","precise":"Inputs $x \\sim N(0, I_d)$ and labels $y = g(W x)$ with $W$ an $r$ x d matrix of rank r fixed as $d \\to \\infty$ and $g$ a fixed link function. For a standard two-layer network trained by online or multi-pass SGD on the square loss without data preprocessing, determine the exponent $\\gamma(g)$ in $n \\sim d^\\gamma$ required to reach test error $o(1)$, and determine whether it equals the benchmark for statistical-query and low-degree algorithms, $n \\sim d^{\\operatorname{max}(1, k_*/2)}$, with $k_*$ the generative exponent for single-index targets and the generative leap exponent for multi-index targets (Damian, Lee and Bruna, arXiv:2506.05500). The answer is $\\gamma$ as a function of the Hermite structure of $g$, with proof.","problem_ref":null,"references":"","settled_by":"Matching upper and lower bounds on the sample complexity of standard gradient training for general multi-index targets.","status_note":"For single-index polynomial targets, SGD with minibatch reuse on two-layer networks reaches $n \\sim d\\,\\operatorname{polylog} d$ (Lee, Oko, Suzuki and Wu, arXiv:2406.01581, 2024); correlational SGD on multi-index targets is governed by the leap complexity (Abbe, Boix-Adsera and Misiakiewicz 2023); the low-degree benchmark for multi-index targets is set by the generative leap (Damian, Lee and Bruna 2025); whether standard SGD reaches that benchmark for general multi-index $g$ is open (2026).","title":"Sample complexity of gradient training for multi-index targets","topic_ref":"e915a7e12556cfb0adf583a894af7480127cf4267dc4e2458650f4872a864ed1"},"attested":{"attestation":{"batch":null,"client_id":null,"id_token_sha256":null,"kind":"platform"},"record_hash":"91fdc52a2137848f1217c031310bc0a08f0dfe8fe5bf8dd82fb22aec4ec71ab9","schema":"pubphys.attested/1"},"envelope":{"attested_hash":"56cdf804493d55834aab980d064e6a0319a55a0c628cbda1bb74a534d9ac72da","platform_signature":{"key_id":"c6afc19b31429869751f06879c75cd64ea92654423d15b44be775bf1310a60da","sig":"YBzXrxTc8bxJ1Y0HDVZ-eMaOQTk3FtIrCY5cAUDzyoWfWUm2qG3dAlO4m1zYYaU5SxkOuN8IM6QuyrVDnswRCw"},"schema":"pubphys.envelope/1"},"record_hash":"91fdc52a2137848f1217c031310bc0a08f0dfe8fe5bf8dd82fb22aec4ec71ab9","leaf_index":2095}