A tiny network memorizes modular addition, then suddenly understands it.
A small network learns (a + b) mod 47 from 45% of all pairs, live. Each number has a learned embedding, the hidden layer squares the sum of two embeddings, and a linear readout scores every possible answer; it trains full batch with AdamW from a large initialization. It memorizes its training pairs within a few dozen steps while test accuracy sits at chance, then hundreds of steps later, as weight decay shrinks the weights, it abruptly generalizes. Underneath it has found a Fourier algorithm: the neuron map shows each hidden unit locking onto a single frequency, and the ring projects the embeddings onto the strongest one, where adding two unseen numbers is just stacking their angles. The table of answers turns from noise into rainbow diagonals.
Try it. Toggle weight decay mid-run, pick how much of the table is used for training, or switch the modulus between 31, 47 and 59. Press W for weight decay, 1 to 3 for the training fraction, P to cycle the prime and R to restart with a new seed.
Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.
Build a live grokking demo with JavaScript and the HTML canvas element: a tiny network learns modular addition, memorizes first, and only generalizes much later. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.
Start simple:
- Make a canvas that fills the window, stays sharp on high-DPI screens (scale by devicePixelRatio), and resizes with the window.
- Pick a prime p around 31 to 47. List the unordered pairs (a, b), shuffle them, and use about half for training (with both orders) and the rest as a test set, so the test set never contains b + a for a trained a + b.
- The model: a learned embedding E (64 numbers per token, initialized fairly large), hidden layer h = (E[a] + E[b]) squared elementwise, and a linear readout W giving one logit per possible answer.
- Train full batch with softmax cross-entropy and AdamW (Adam with decoupled weight decay around 1.0 and a learning rate near 0.01) as many steps per frame as fit in about 8 ms.
- Every few steps, measure train and test accuracy and plot both against the step count on a log scale. Train accuracy should hit 100% fast while test stays at chance, then jumps much later.
Once that works, make it beautiful:
- Draw the model's answer for every pair as a p by p image colored by hue. It turns from noise into clean rainbow diagonals when the net generalizes.
- Take the discrete Fourier transform of each embedding row over a, find the strongest frequency k, and project the embeddings onto its cosine and sine directions. Draw the numbers as beads: a cloud, then a ring.
- Add a weight decay toggle so I can see that without it the network never generalizes.
Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as a neuron by frequency heat map, animating addition as rotation on the ring, or comparing training fractions.