Optimization landscape of two-layer neural networks (population risk)
Characterize the optimization landscape of the population risk R_N(θ) for two-layer neural networks with prediction ŷ(x;θ) = (1/N) ∑_{i=1}^N σ_*(x; θ_i) and square loss ℓ(y,ŷ) = (y − ŷ)^2, including the existence and structure of local minima, saddle points, and global minima, even when an infinite number of training examples are available.
References
Understanding the optimization landscape of two-layers neural networks is largely an open problem even when we have access to an infinite number of examples, i.e. to the population risk R_{N}(\theta).
\begin{conjecture} \label{conj:higher-rank} Let $d\ge1$, let the teacher network satisfy $(t,v)\inP{d}{m}$ with any linear skip $b{\mathbf t}\inRd$ and let $n\ge m$. Then every local minimum $(b{\mathbf s},s,w)$ of $L$ on $Rd\timesRn\times(S{d-1})n$ with $(s,w)\inP{d}{n}$ is an exact fit. \end{conjecture}
Like most of the loss landscape theory it builds on, our analysis is confined to one hidden layer: the residual representation of \Cref{prop:loss_in_residual} rests on the closed form of the pair moments, which has no counterpart for deeper networks, so the effect of a learned skip on the landscape of deep networks remains open.
Whether \Cref{prop:nc-dead-student} holds for $d=2$ and non-negative teacher masses, as \Cref{prop:zero_mass_exact_fit} does for the centered loss, is open.