A tiny transformer learns to braid, reverse and sort digits, attention woven as thread.
A one-layer, four-head causal transformer trains in your browser with hand-written backpropagation and Adam, on fresh random sequences of eight digits followed by a separator and the answer, predicting each answer digit like the next token. To test it, the model generates its output greedily, feeding each guess back in, and the attention weights of that pass become threads from each output spool back to the input spools, one natural dye per head, with thickness and tension set by the weight. Untrained attention is spread thin, so the threads hang slack and tangled; as the loss falls they pull taut into braids, crossings or the data-dependent permutation of a sort. Below the loom, the attention matrices themselves are woven into a kilim, one row per shuttle pass, so the cloth records the moment the model figured the task out.
Try it. Type digits to feed your own sequence, or click and drag a top spool to change one digit. Pick a task (braid, reverse, sort, riffle or rotate) to train a fresh model, click a head to isolate its threads and its stripe of cloth, and hover an output to see only its threads. T or the arrow keys change task, space pauses training, R relearns from scratch.
Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.
Build a tiny transformer that learns to reverse a sequence of digits right in the browser, and draw its attention as colored threads. Use JavaScript and the HTML canvas element, in a single index.html file with no libraries or build step, so I can open it directly in a browser.
Start simple:
- Make a canvas that fills the window, stays sharp on high-DPI screens (scale by devicePixelRatio), and resizes with the window. Paint it a dark background.
- Write a minimal single-head attention model with no libraries: token embeddings for the digits 0 to 9, learned position embeddings for 6 output slots and 6 input slots, and query, key and value matrices. Each output slot's query attends over the 6 input positions with a softmax, and the weighted sum of values goes through a final matrix to 10 digit logits.
- Train it to reverse random 6-digit sequences with cross-entropy loss. Write the backward pass by hand (the softmax gradient is a * (da - sum(a * da))) and use plain gradient descent or Adam on fresh random examples, a few steps per animation frame.
- Draw the input digits in a row at the top and the predicted digits at the bottom. For every output and input pair, draw a curve between them whose thickness and opacity are the attention weight. Show the loss as a number.
Once that works, make it beautiful:
- Add a causal decoder version: train on 'input, separator, answer' sequences with masked self-attention, and generate the answer one digit at a time, feeding each guess back in.
- Use several heads, each a different dye color, and offset their threads so they lie side by side like yarn. Let weak threads sag and sway while strong ones run straight.
- Let me type digits to test it, and add buttons for other tasks like sort or swapping adjacent pairs.
Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as weaving the attention matrices into a scrolling tapestry, adding a second layer to learn harder tasks, or plotting the loss curve.