Visualizations
423 / 500

423 · Machine learning

Latent Face Atlas

A VAE learns cartoon faces in seconds and lays out its latent space as a map of faces.

A variational autoencoder trains in your browser on 600 cartoon faces drawn procedurally from two hidden factors, mood (a worried frown through to a wide, rosy smile) and build (a narrow face under a low fringe through to a wide round one), which the model is never told. The encoder squeezes each 24 x 24 face into the mean and variance of a 2D code, a code is sampled with the reparameterization trick, and the decoder paints the face back, trained on per-pixel cross-entropy plus a KL penalty that is eased in so the second latent dimension is not abandoned; backprop and Adam are hand-written over Float32Arrays. The map decodes a grid of codes into a sheet of faces that morph smoothly from frown to smile and narrow to wide, upsampling the decoded gray levels before colorizing them so every face keeps crisp edges, and each training face sits at its encoded position as a dot colored by its true mood, so you can watch the model discover the factors on its own. Once the VAE has a head start, a plain autoencoder of the same shape trains alongside for comparison: its codes huddle into a thin filament wherever they like, and the faces decoded away from it dissolve into speckle.

Try it. Move or drag across the map to decode any point in latent space into the large face on the right (on a phone, it floats beside your finger), or nudge the cursor with the arrow keys. Switch between the VAE and the plain autoencoder with the chips or A (the autopilot flips between them every so often), Space pauses training and R retrains both from scratch.

  • Variational autoencoder
  • Reparameterization trick
  • KL warm-up
  • Backprop and Adam from scratch
  • Contour-preserving upsampling

View the source · one module, plus a small shared runtime for sizing, the animation loop and input

Build your own

Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.

Build a variational autoencoder (VAE) that trains in the browser on cartoon faces and shows its 2D latent space as a map of faces, using JavaScript and the HTML canvas element. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.

Start simple:
- Generate a dataset of 500 grayscale 24 x 24 cartoon faces in code from two random factors in [-1, 1]: mood (the mouth goes from a frown to an open smile, and the eyebrows tilt) and build (the face ellipse gets wider and the hairline moves). Use a few gray levels: 0 background, 0.55 skin, 0.8 hair, 1 for eyes, brows and mouth, and supersample each pixel 3 x 3 for smooth edges.
- Write a tiny multilayer perceptron by hand with Float32Arrays, with backpropagation and the Adam optimizer. Make an encoder (576 -> 32 -> 4: a 2D mean and a 2D log-variance) and a decoder (2 -> 32 -> 32 -> 576 logits).
- Train on batches of 16: sample z = mean + exp(0.5 * logvar) * eps, decode, and minimize per-pixel binary cross-entropy plus the KL divergence to a unit Gaussian. Ramp the KL weight from 0 to 1 over the first 1,000 steps so both latent dimensions get used. Run steps every frame within about 8 ms.
- Draw an 11 x 11 grid of faces decoded from z values between -2.2 and 2.2, refreshing a row or two per frame.

Once that works, make it beautiful:
- Color the gray levels with a lookup table (dark navy background, peach skin, auburn hair, dark ink) so the faces look like stickers.
- Plot every training face at its encoded mean as a small dot colored by its true mood, and watch the clusters organize.
- Show a large decoded face under the mouse as I move across the map.

Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as training a plain autoencoder side by side to see the holes in its latent space, adding a third factor and a 3D latent, or interpolating between two faces I pick.
PreviousSnow and SlushSnowballs, jelly, sand and water smash and tumble down a slope, simulated with MLS-MPM. NextKolamA hand traces a rice-flour kolam: one unbroken line looping around every dot.

Related visualizations

  • GAN DuelMachine learning A generator and a discriminator fight live over a 2D target, failure modes and all.
  • Normalizing FlowMachine learning An invertible network trains live and bends a Gaussian and its grid like taffy.
  • Diffusion SketchpadMachine learning A diffusion model trains live and condenses pure noise into whatever shape you draw.
  • Attention LoomMachine learning A tiny transformer learns to braid, reverse and sort digits, attention woven as thread.
  • GrokkingMachine learning A tiny network memorizes modular addition, then suddenly understands it.
  • Fourier FeaturesMachine learning A plain MLP, random Fourier features and a SIREN race to memorize one picture.
  • Skip-gram ConstellationMachine learning Word2vec learns from scratch, and king minus man plus woman lands on queen.
  • Gaussian SplattingMachine learning 3D Gaussian splatting in a 2D canvas, trained from scratch by gradient descent.
  • ReLU Stained GlassMachine learning A ReLU network's exact linear regions, drawn as a stained glass window.

Use ← and → to move between demos. While the canvas has focus, keys go to the demo instead.

← More from Emergent Mind Labs