Prune a network to a sparse winning ticket that still learns, round by round.
The lottery ticket hypothesis, reproduced live. A ReLU network with two hidden layers of 32 learns a two-arm spiral, then iterative magnitude pruning runs: cut the smallest surviving weights, rewind the rest to the values they had at step 30 of the first run, and train again. The survivors form a winning ticket that keeps learning the spiral with only a few percent of its weights, while a random ticket with exactly as many weights per layer and the same starting values falls apart much sooner. Every weight is drawn as a thread colored by sign; pruned threads burn out and the survivors visibly morph back to their early values, and each round adds a point per ticket to the accuracy versus sparsity plot.
Try it. Pick how much to prune per round (20, 40 or 60 percent, or keys 1 to 3), cut early with Prune now or the space bar, and restart with a new seed with R.
Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.
Build a live lottery ticket hypothesis demo with JavaScript and the HTML canvas element: prune a small network round by round and show that a sparse "winning ticket" still learns. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.
Start simple:
- Make a canvas that fills the window, stays sharp on high-DPI screens (scale by devicePixelRatio), and resizes with the window.
- Generate a two-arm spiral dataset (a training set and a separate test set) in the square from -1 to 1.
- Write a ReLU network from scratch: 2 inputs, two hidden layers of 32 units, one logit. Keep a 0/1 mask per weight. Train with full-batch Adam and binary cross-entropy, keeping masked weights at zero.
- Save a copy of the weights at step 30 of the first run (the rewind point).
- Run rounds: train 400 steps, record test accuracy, prune the 40% smallest-magnitude surviving weights in each layer, reset the survivors to their step-30 values, and train again.
- As a control, each round also train a random ticket: a random mask with the same number of weights per layer, starting from the same step-30 values.
Once that works, make it beautiful:
- Draw the network: neurons as dots in columns and every surviving weight as a curved thread, warm for positive and cool for negative, thicker for larger magnitude, with additive blending.
- Animate each prune by flashing the doomed threads white-hot before they vanish, and the rewind by morphing the survivors back to their early values.
- Plot test accuracy against the fraction of weights remaining on a log axis, one line for the winning ticket and one for the random ticket, plus both decision boundaries.
Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as comparing rewinding to step 0, trying global instead of per-layer pruning, or highlighting neurons that lose every connection.