Optimization landscape of two-layer neural networks (population risk)

Characterize the optimization landscape of the population risk R_N(θ) for two-layer neural networks with prediction ŷ(x;θ) = (1/N) ∑_{i=1}^N σ_*(x; θ_i) and square loss ℓ(y,ŷ) = (y − ŷ)^2, including the existence and structure of local minima, saddle points, and global minima, even when an infinite number of training examples are available.

Background

Two-layer neural networks induce a high-dimensional non-convex risk landscape, and while specific models have been analyzed, a general understanding of this landscape remains elusive. Even in the infinite-sample (population risk) setting, where stochastic fluctuations due to finite datasets are absent, the structure of local and global minima of R_N(θ) is not comprehensively understood.

This paper introduces a mean-field (distributional dynamics) perspective via a PDE to analyze training dynamics and generalization, but the full characterization of the population risk landscape for two-layer networks is identified as an outstanding problem.

References

Understanding the optimization landscape of two-layers neural networks is largely an open problem even when we have access to an infinite number of examples, i.e. to the population risk R_{N}(\theta).

— A Mean Field View of the Landscape of Two-Layers Neural Networks  (1804.06561 - Mei et al., 2018) in Introduction

\begin{conjecture} \label{conj:higher-rank} Let $d\ge1$, let the teacher network satisfy $(t,v)\inP{d}{m}$ with any linear skip $b{\mathbf t}\inRd$ and let $n\ge m$. Then every local minimum $(b{\mathbf s},s,w)$ of $L$ on $Rd\timesRn\times(S{d-1})n$ with $(s,w)\inP{d}{n}$ is an exact fit. \end{conjecture}

— Removing spurious minima for planar features by skip connections  (2610.01728 - Zimmermann et al., 1 Oct 2026) in Section 6, Conclusion and Limitations; Conjecture (labelled conj:higher-rank)

Like most of the loss landscape theory it builds on, our analysis is confined to one hidden layer: the residual representation of \Cref{prop:loss_in_residual} rests on the closed form of the pair moments, which has no counterpart for deeper networks, so the effect of a learned skip on the landscape of deep networks remains open.

— Removing spurious minima for planar features by skip connections  (2610.01728 - Zimmermann et al., 1 Oct 2026) in Section 6, Conclusion and Limitations

Whether \Cref{prop:nc-dead-student} holds for $d=2$ and non-negative teacher masses, as \Cref{prop:zero_mass_exact_fit} does for the centered loss, is open.

— Removing spurious minima for planar features by skip connections  (2610.01728 - Zimmermann et al., 1 Oct 2026) in Appendix, subsection “Benign Loss Landscape for Coplanar Teacher Networks,” immediately after the proof of Theorem thm:centered