Conference item icon

Conference item

Convergence analysis of Newton's method for neural networks in the overparameterized limit

Abstract:
A convergence analysis is developed for the regularized Newton method for training neural networks (NNs) in the overparameterized limit. We prove that as the number of hidden units tends to infinity, the NN training dynamics converge in probability to the solution of a deterministic limit equation involving a “Newton neural tangent kernel” (NNTK). Explicit rates characterizing this convergence are provided and, in the infinite-width limit, we prove that the NN converges exponentially fast to the target data (i.e., a global minimizer with zero loss). Crucially, we show that this convergence is uniform across the frequency spectrum, addressing the spectral bias inherent in gradient descent. The eigenvalues of the NTK for gradient descent accumulate at zero, leading to slow convergence for target data with high-frequency components. In contrast, the NNTK has uniformly lower bounded eigenvalues if the regularization parameter is selected appropriately, allowing Newton’s method to converge more quickly for data with high-frequency components. Mathematical challenges that our analysis needs to address include the implicit parameter update of the Newton method with a potentially indefinite Hessian matrix and the fact that the dimension of this linear system of equations tends to infinity as the NN width grows. This substantially complicates deriving the training dynamics in the overparameterized limit as well as proving the convergence of the finite-width dynamics thereto. Our analysis identifies a scaling formula for selecting the regularization parameter, which we show can vanish at a suitable rate as the NN width becomes larger. In addition, we prove that, for sufficiently large numbers of hidden units, the regularized Hessian remains positive definite during training and the Newton updates for individual NN parameters converge to zero, demonstrating that the model behaves as a linearization around the initialization.
Publication status:
Accepted
Peer review status:
Peer reviewed

Actions

Authors

More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Mathematical Institute
Role:
Author
ORCID:
0000-0002-2206-4334
More by this author
Institution:
University of Oxford
Division:
MPLS
Department:
Mathematical Institute
Role:
Author


Publisher:
NeurIPS
Acceptance date:
2026-09-24
Event title:
40th Annual Conference on Neural Information Processing Systems (NeurIPS 2026)
Event location:
Sydney, Australia
Event website:
https://neurips.cc/Conferences/2026
Event start date:
2026-12-06
Event end date:
2026-12-12


Language:
English
Pubs id:
2460233
Local pid:
pubs:2460233
Deposit date:
2026-09-25
ARK identifier:

Terms of use


Views and Downloads

Views and downloads will return soon






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP