Visualizations
384 / 500

384 · Algorithms

Cache Lines

Four loop orders of one matrix multiply race through a simulated cache hierarchy.

The same 32 by 32 matrix multiply runs in four loop orders at once, each against its own simulated memory: a set-associative L1 with LRU replacement, an 8 KB L2 and RAM, charged an estimated 1, 5 and 80 nanoseconds. Every access to A, B and C looks up the real 64-byte line that holds the element. All four orders get the same simulated time per frame, so the ones that stream along rows finish first while j-k-i, which walks down columns of A and C, takes about ten times longer. In the detail view every line of every matrix is lit by where it lives right now, misses fly up from RAM, and C fills in with the color of when each sum finished, revealing the traversal pattern.

Try it. Click a race lane (or press 1 to 4) to show that order in detail. Change the L1 size (L) and its associativity (W) and watch conflict misses appear or vanish, or slow the clock to 1/500 (S) to follow single accesses. Hover any matrix element to see its cache line and the one set it can live in.

  • Set-associative LRU cache simulation
  • Loop tiling
  • Conflict misses
  • Completion-order heatmaps

View the source · one module, plus a small shared runtime for sizing, the animation loop and input

Build your own

Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.

Build an interactive visualization that shows how loop order changes cache behavior in a matrix multiply, using JavaScript and the HTML canvas element. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.

Start simple:
- Make a canvas that fills the window, stays sharp on high-DPI screens (scale by devicePixelRatio), and resizes with the window. Use a dark background.
- Simulate C += A x B for 32 by 32 matrices of 8-byte numbers stored row by row, with A, B and C placed one after another in memory. You only need addresses, not real values.
- Write a small cache simulator: 64-byte lines (8 numbers each), a fixed number of sets, a few ways per set and least recently used replacement. An access is a hit if its line is in the cache; otherwise load it, evicting the oldest line in that set.
- Run the triple loop in i-j-k order, a few hundred iterations per frame, touching A[i][k], B[k][j] and C[i][j]. Count hits and misses and add up an estimated time: 1 ns per hit and 80 ns per miss.
- Draw the three matrices as grids, lighting each element whose line is currently in the cache, and outline the elements touched by the current iteration.

Once that works, make it beautiful:
- Draw each cache line as one rounded bar of 8 cells so you can see whole lines arrive and leave.
- Run i-j-k, i-k-j, j-k-i and a tiled version side by side, each with its own cache, giving every order the same simulated time per frame, and show progress bars so the race is visible.
- Show the cache itself as a grid of sets and ways, colored by which matrix each line belongs to, and flash slots on hits and misses.

Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as adding an L2 cache, sliders for cache size and associativity, or coloring C by when each element finished.
PreviousAperture ProblemOne grating, one true motion, three holes that each report a different direction. NextTessellation MorphDrag one edge of an Escher-style creature and every copy reshapes while still tiling.

Related visualizations

  • Core ArenaAlgorithms Assembly warriors battle for 8,000 cells of shared memory on a real Redcode VM.
  • Marching CubesAlgorithms Watch the classic isosurface algorithm march cell by cell, then zoom out to clay.
  • CPU PipelineRetro A five-stage RISC pipeline as an isometric factory: bubbles, bypasses and flushes.
  • Mode 7 TrackRetro The per-scanline floor trick of 16-bit kart racers, with a debug view of every row.
  • Spatial IndexesAlgorithms Quadtree, k-d tree, BVH and grid rebuilt every frame over thousands of moving points.
  • Breadboard CPURetro A microcoded 8-bit breadboard computer paints cellular automata on a 16x16 LED matrix.
  • Dancing LinksAlgorithms Knuth's Algorithm X tiles boards with wooden pentominoes while its links visibly dance.
  • Sorting TapestryAlgorithms Six sorting algorithms cross-stitch a scrambled sampler back into a rainbow.
  • Self-Replicating LoopsEmergence Langton's 1984 loops copy their own genome and tile the plane with a glass colony.

Use ← and → to move between demos. While the canvas has focus, keys go to the demo instead.

← More from Emergent Mind Labs