Visualizations
421 / 500

421 · Type and image

BPE Tokenizer

Byte pair encoding learns a vocabulary live, fusing typeset tiles merge by merge.

Training starts with every character of the opening of A Tale of Two Cities as its own token. The text is cut into chunks the way GPT-2 does it (an optional leading space plus a run of letters), then the most frequent adjacent pair of tokens is found, added to the vocabulary and replaced everywhere, again and again. Before each merge the pair lights up across the passage, then the tiles slide together into one new color, and tiles keep their identity through every change so the layout reflows smoothly. The side panel lists the latest merges with their counts, draws the derivation tree of the selected token and plots the token count falling, and a second box re-encodes any text by replaying the learned merges in order.

Try it. Click and type to tokenize your own text with the current vocabulary (Backspace deletes, Esc clears). Drag along the chart to scrub through the merge history, use the left and right arrow keys to step one merge, and hover or click a tile to see where it came from. Left alone, it trains to the end, demonstrates on new sentences and starts over.

  • Byte pair encoding training and encoding
  • GPT-2 style pre-tokenization
  • Identity-preserving tile layout animation
  • Cached measureText flow layout

View the source · one module, plus a small shared runtime for sizing, the animation loop and input

Build your own

Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.

Build an animated byte pair encoding (BPE) tokenizer with JavaScript and the HTML canvas element. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.

Start simple:
- Make a canvas that fills the window, stays sharp on high-DPI screens (scale by devicePixelRatio), and resizes with the window. Paint it a dark background.
- Put a public domain passage in a string, for example the opening of A Tale of Two Cities, which repeats itself a lot.
- Split it into chunks with a regex like / ?[A-Za-z]+| ?[^\sA-Za-z]+|\s+/g so a word keeps its leading space and merges never cross words. Start every chunk as a list of single characters.
- Training step: count every adjacent pair of tokens across all chunks, take the most frequent pair, record it as a merge, and replace it everywhere with the joined string. Stop when no pair appears twice.
- Draw the passage as a row-wrapped sequence of tiles, one per token: a rounded rectangle sized with ctx.measureText (cache the widths) and the token text inside. Show spaces as small dots.
- Run one merge every half second and redraw, with a counter showing the token count falling.

Once that works, make it beautiful:
- Give each learned token its own color (golden angle hues) and keep single characters grey, so the passage colors in as the vocabulary grows.
- Animate: before a merge, outline every occurrence of the pair; after it, start the new tile at the combined position of the old two and ease every tile to its new place.
- Add a list of recent merges and a small chart of token count versus merges that you can drag to scrub back and forth.
- Add a second box where I can type, encoded by replaying the learned merges in order.

Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as working on raw UTF-8 bytes, drawing each token's merge tree, or comparing vocabularies learned from two different texts.
PreviousRogue WavesA storm swell where walls of water rise from nowhere through nonlinear focusing. NextSnow and SlushSnowballs, jelly, sand and water smash and tumble down a slope, simulated with MLS-MPM.

Related visualizations

  • Compression EchoesType and image LZ77 compresses a poem while every back-reference loops back over the page in gold.
  • Huffman GardenAlgorithms Letter frequencies sprout into a Huffman tree drawn as a botanical plate.
  • Temperature GardenMachine learning Greedy, beam, top-k, nucleus and temperature sampling grown as a botanical plate.
  • Particle TypeType and image Words made of thousands of particles that scatter and regroup.
  • Bytebeat MachineSound Whole songs from one line of bit arithmetic, drawn as a tapestry of bytes.
  • Hidden in PixelsType and image A treasure map hides in a landscape's lowest bits. Peel the bit planes to find it.
  • Multiscale TruchetGenerative art Truchet tiles at five scales lock together seamlessly, flipping color at each level.
  • Line BreakingType and image Greedy line breaking versus Knuth-Plass, with rivers and the breakpoint graph.
  • Skip-gram ConstellationMachine learning Word2vec learns from scratch, and king minus man plus woman lands on queen.

Use ← and → to move between demos. While the canvas has focus, keys go to the demo instead.

← More from Emergent Mind Labs