Extend the μP learning-rate transfer proof to nonlinear MLPs and other optimizers
Establish a rigorous proof of learning-rate transfer with width under Maximal Update Parametrization (μP) for non-linear multi-layer perceptrons with activation functions (e.g., ReLU) and for training algorithms beyond gradient descent (e.g., Adam), including the development of proof techniques to control large-width deviations in these settings.
References
While our results are limited to linear networks trained with GD, we believe they can be extended to non-linear MLPs and different optimizers. However, this will likely require different proof machinery especially when dealing when large-width deviations. We leave this question for future work.
Finally, two formal-theory questions remain open: whether the two-condition decomposition is jointly sufficient for transfer at infinite width, and a faithful test of $P$-SSM under ZOH SSMs with growing $N_x$, the regime in which \citet{vankadara2024feature}'s prescription was derived (our comparisons apply it to the field-standard simplified-ZOH Mamba~\citep{gu2023mamba} at fixed $N_x{=}16$; Appendix~\ref{sec:appendix:vankadara_faithful}).