High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks
Simon Martin, Giulio Biroli, Francis Bach; 27(136):1−182, 2026.
Abstract
We study the high-dimensional training dynamics of a shallow neural network with quadratic activation in a teacher--student setup. We focus on the extensive-width regime, where the teacher and student network widths scale proportionally with the input dimension, and the sample size grows quadratically. This scaling aims to describe overparameterized neural networks in which feature learning still plays a central role. In the high-dimensional limit, we derive a dynamical characterization of the gradient flow, in the spirit of dynamical mean-field theory (DMFT). Under $\ell_2$-regularization, we analyze these equations at long times and characterize the performance and spectral properties of the resulting estimator. This result provides a quantitative understanding of the effect of overparameterization on learning and generalization, and reveals a double descent phenomenon in the presence of label noise, where generalization improves beyond interpolation. In the small regularization limit, we obtain an exact expression for the perfect recovery threshold as a function of the network widths, providing a precise characterization of how overparameterization influences recovery.
[abs]
[pdf][bib] [code]| © JMLR 2026. (edit, beta) |
Facts Only
* The study concerns training dynamics of a shallow neural network with quadratic activation in a teacher-student setup.
* Focus is placed on the extensive-width regime where teacher and student widths scale proportionally to the input dimension, and sample size grows quadratically.
* Dynamical characterization of gradient flow is derived in the high-dimensional limit, inspired by dynamical mean-field theory (DMFT).
* Analysis is performed under $\ell2$-regularization at long times to characterize estimator performance and spectral properties.
* The results reveal a double descent phenomenon in the presence of label noise, where generalization surpasses interpolation.
* An exact expression for the perfect recovery threshold is obtained in the small regularization limit as a function of network widths.
Executive Summary
Full Take
The work connects high-dimensional scaling regimes with fundamental learning phenomena like generalization and recovery thresholds. The emergence of double descent in this context suggests that excessive capacity does not merely lead to perfect interpolation but introduces a phase transition where optimization landscapes favor better generalization, even when noise is introduced. This implies that the benefits of overparameterization are contingent on the specific structure of the training dynamics—specifically, how gradients flow in high dimensions. The exact expression for the recovery threshold links structural properties (network widths) directly to the attainable performance limits under regularization constraints. The connection between gradient flow dynamics and generalization points toward a potential mechanism where dynamical stability dictates the quality of learned representations beyond simple asymptotic results. The focus shifts from static generalization bounds to the dynamic process by which knowledge is acquired and refined in massively overparameterized systems.
Bridge Questions: What are the necessary and sufficient conditions on the network width scaling for the double descent phenomenon to persist under varying noise levels? How does the dynamical characterization of gradient flow provide insight into the mechanism causing the shift from interpolation to generalization performance? If the exact recovery threshold is known, how do these theoretical limits inform the design of regularization techniques for overparameterized models in real-world scenarios?
Sentinel — Human
The text exhibits the dense, precise structure of a genuine academic research abstract, suggesting human authorship focused on conveying specific theoretical results.
