Byte pair encoding learns a vocabulary live, fusing typeset tiles merge by merge.
Training starts with every character of the opening of A Tale of Two Cities as its own token. The text is cut into chunks the way GPT-2 does it (an optional leading space plus a run of letters), then the most frequent adjacent pair of tokens is found, added to the vocabulary and replaced everywhere, again and again. Before each merge the pair lights up across the passage, then the tiles slide together into one new color, and tiles keep their identity through every change so the layout reflows smoothly. The side panel lists the latest merges with their counts, draws the derivation tree of the selected token and plots the token count falling, and a second box re-encodes any text by replaying the learned merges in order.
Try it. Click and type to tokenize your own text with the current vocabulary (Backspace deletes, Esc clears). Drag along the chart to scrub through the merge history, use the left and right arrow keys to step one merge, and hover or click a tile to see where it came from. Left alone, it trains to the end, demonstrates on new sentences and starts over.
Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.
Build an animated byte pair encoding (BPE) tokenizer with JavaScript and the HTML canvas element. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.
Start simple:
- Make a canvas that fills the window, stays sharp on high-DPI screens (scale by devicePixelRatio), and resizes with the window. Paint it a dark background.
- Put a public domain passage in a string, for example the opening of A Tale of Two Cities, which repeats itself a lot.
- Split it into chunks with a regex like / ?[A-Za-z]+| ?[^\sA-Za-z]+|\s+/g so a word keeps its leading space and merges never cross words. Start every chunk as a list of single characters.
- Training step: count every adjacent pair of tokens across all chunks, take the most frequent pair, record it as a merge, and replace it everywhere with the joined string. Stop when no pair appears twice.
- Draw the passage as a row-wrapped sequence of tiles, one per token: a rounded rectangle sized with ctx.measureText (cache the widths) and the token text inside. Show spaces as small dots.
- Run one merge every half second and redraw, with a counter showing the token count falling.
Once that works, make it beautiful:
- Give each learned token its own color (golden angle hues) and keep single characters grey, so the passage colors in as the vocabulary grows.
- Animate: before a merge, outline every occurrence of the pair; after it, start the new tile at the combined position of the old two and ease every tile to its new place.
- Add a list of recent merges and a small chart of token count versus merges that you can drag to scrub back and forth.
- Add a second box where I can type, encoded by replaying the learned merges in order.
Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as working on raw UTF-8 bytes, drawing each token's merge tree, or comparing vocabularies learned from two different texts.