Papers
Topics
Authors
Recent
Search
2000 character limit reached

Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

Published 9 Sep 2026 in cs.AI and cs.LG | (2609.10657v1)

Abstract: Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: TgrokH<sup>0.27</sup>D<sup>2.04</sup>η<sup>0.50</sup>λ<sup>0.64T_{\mathrm{grok}} \propto H<sup>{-0.27}\,</sup> D<sup>{-2.04}\,</sup> η<sup>{-0.50}\,</sup> λ<sup>{-0.64} (R<sup>2</sup>=0.732R<sup>2</sup> = 0.732; $0.821$ with interactions). The exponent hierarchy reveals that data complexity (D<sup>2.04D<sup>{-2.04}) is the dominant driver of regime transition, not model capacity (H<sup>0.27H<sup>{-0.27}): doubling data accelerates generalization by 4×{\sim}4\times, while doubling width yields only 1.2×{\sim}1.2\times. A sharp phase boundary at weight decay λ1.0λ\gtrsim 1.0 separates grokking from non-grokking configurations, and weight norm trajectories show monotonic compression during the transition, consistent with implicit regularization selecting low-complexity solutions. These results provide a quantitative foundation for predicting and controlling regime transitions in overparameterized networks.

Authors (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.