Visualizations
144 / 500

144 · Machine learning

Grokking

A tiny network memorizes modular addition, then suddenly understands it.

A small network learns (a + b) mod 47 from 45% of all pairs, live. Each number has a learned embedding, the hidden layer squares the sum of two embeddings, and a linear readout scores every possible answer; it trains full batch with AdamW from a large initialization. It memorizes its training pairs within a few dozen steps while test accuracy sits at chance, then hundreds of steps later, as weight decay shrinks the weights, it abruptly generalizes. Underneath it has found a Fourier algorithm: the neuron map shows each hidden unit locking onto a single frequency, and the ring projects the embeddings onto the strongest one, where adding two unseen numbers is just stacking their angles. The table of answers turns from noise into rainbow diagonals.

Try it. Toggle weight decay mid-run, pick how much of the table is used for training, or switch the modulus between 31, 47 and 59. Press W for weight decay, 1 to 3 for the training fraction, P to cycle the prime and R to restart with a new seed.

  • AdamW with weight decay
  • Discrete Fourier analysis
  • Cross-entropy training

View the source · one module, plus a small shared runtime for sizing, the animation loop and input

Build your own

Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.

Build a live grokking demo with JavaScript and the HTML canvas element: a tiny network learns modular addition, memorizes first, and only generalizes much later. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.

Start simple:
- Make a canvas that fills the window, stays sharp on high-DPI screens (scale by devicePixelRatio), and resizes with the window.
- Pick a prime p around 31 to 47. List the unordered pairs (a, b), shuffle them, and use about half for training (with both orders) and the rest as a test set, so the test set never contains b + a for a trained a + b.
- The model: a learned embedding E (64 numbers per token, initialized fairly large), hidden layer h = (E[a] + E[b]) squared elementwise, and a linear readout W giving one logit per possible answer.
- Train full batch with softmax cross-entropy and AdamW (Adam with decoupled weight decay around 1.0 and a learning rate near 0.01) as many steps per frame as fit in about 8 ms.
- Every few steps, measure train and test accuracy and plot both against the step count on a log scale. Train accuracy should hit 100% fast while test stays at chance, then jumps much later.

Once that works, make it beautiful:
- Draw the model's answer for every pair as a p by p image colored by hue. It turns from noise into clean rainbow diagonals when the net generalizes.
- Take the discrete Fourier transform of each embedding row over a, find the strongest frequency k, and project the embeddings onto its cosine and sine directions. Draw the numbers as beads: a cloud, then a ring.
- Add a weight decay toggle so I can see that without it the network never generalizes.

Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as a neuron by frequency heat map, animating addition as rotation on the ring, or comparing training fractions.
PreviousLagrange LandscapeThe restricted three-body problem as a plaster relief, Trojans looping on its summits. NextBirdsong SyrinxBirdsong from a two-knob nonlinear oscillator, printed live as an ink sonogram.

Related visualizations

  • Fourier FeaturesMachine learning A plain MLP, random Fourier features and a SIREN race to memorize one picture.
  • Growing Neural CellsMachine learning Cells running one tiny learned rule grow a gecko from a single seed and heal its wounds.
  • Neural NetworkMachine learning A tiny neural network learns to classify points, live, in your browser.
  • Diffusion SketchpadMachine learning A diffusion model trains live and condenses pure noise into whatever shape you draw.
  • Attention LoomMachine learning A tiny transformer learns to braid, reverse and sort digits, attention woven as thread.
  • Lottery TicketMachine learning Prune a network to a sparse winning ticket that still learns, round by round.
  • Normalizing FlowMachine learning An invertible network trains live and bends a Gaussian and its grid like taffy.
  • Hopfield MemoryMachine learning A lamp panel of 400 neurons that repairs scribbled glyphs, until it stores too many.
  • Spiking RasterMachine learning A thousand Izhikevich neurons learn by spike timing and ripple with travelling waves.

Use ← and → to move between demos. While the canvas has focus, keys go to the demo instead.

← More from Emergent Mind Labs