Four loop orders of one matrix multiply race through a simulated cache hierarchy.
The same 32 by 32 matrix multiply runs in four loop orders at once, each against its own simulated memory: a set-associative L1 with LRU replacement, an 8 KB L2 and RAM, charged an estimated 1, 5 and 80 nanoseconds. Every access to A, B and C looks up the real 64-byte line that holds the element. All four orders get the same simulated time per frame, so the ones that stream along rows finish first while j-k-i, which walks down columns of A and C, takes about ten times longer. In the detail view every line of every matrix is lit by where it lives right now, misses fly up from RAM, and C fills in with the color of when each sum finished, revealing the traversal pattern.
Try it. Click a race lane (or press 1 to 4) to show that order in detail. Change the L1 size (L) and its associativity (W) and watch conflict misses appear or vanish, or slow the clock to 1/500 (S) to follow single accesses. Hover any matrix element to see its cache line and the one set it can live in.
Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.
Build an interactive visualization that shows how loop order changes cache behavior in a matrix multiply, using JavaScript and the HTML canvas element. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.
Start simple:
- Make a canvas that fills the window, stays sharp on high-DPI screens (scale by devicePixelRatio), and resizes with the window. Use a dark background.
- Simulate C += A x B for 32 by 32 matrices of 8-byte numbers stored row by row, with A, B and C placed one after another in memory. You only need addresses, not real values.
- Write a small cache simulator: 64-byte lines (8 numbers each), a fixed number of sets, a few ways per set and least recently used replacement. An access is a hit if its line is in the cache; otherwise load it, evicting the oldest line in that set.
- Run the triple loop in i-j-k order, a few hundred iterations per frame, touching A[i][k], B[k][j] and C[i][j]. Count hits and misses and add up an estimated time: 1 ns per hit and 80 ns per miss.
- Draw the three matrices as grids, lighting each element whose line is currently in the cache, and outline the elements touched by the current iteration.
Once that works, make it beautiful:
- Draw each cache line as one rounded bar of 8 cells so you can see whole lines arrive and leave.
- Run i-j-k, i-k-j, j-k-i and a tiled version side by side, each with its own cache, giving every order the same simulated time per frame, and show progress bars so the race is visible.
- Show the cache itself as a grid of sets and ways, colored by which matrix each line belongs to, and flash slots on hits and misses.
Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as adding an L2 cache, sliders for cache size and associativity, or coloring C by when each element finished.