Visualizations
168 / 500

168 · Machine learning

Deep Q Snake

A deep Q-network learns snake from scratch, live, on a retro dot-matrix LCD.

A small neural network (17 inputs, two hidden layers, 3 outputs) learns to play snake by deep Q-learning, starting from random weights every time the page loads. It sees the board from the snake's point of view: for each of its three moves, whether that square is deadly, how far it can see, and how much of the board a flood fill can still reach from there, plus where the food is. Every move goes into a 40,000-slot replay memory, and each step a random minibatch of 32 is replayed to pull Q(s, a) toward the reward plus the discounted value of the next state, using Double DQN targets from a frozen target network that is refreshed every 250 updates. A hidden training snake takes as many steps as fit in a few milliseconds each frame, so within seconds the average score climbs from zero into the twenties or thirties, while the snake on the LCD plays with the same network and its arrows show the Q-value of turning left, going ahead or turning right.

Try it. Take over with the arrow keys, WASD or a swipe (your moves are learned from too, and the agent's arrows still show what it would do); it hands back to the agent when you stop. Tap Training to switch between Slow, Fast and Turbo, Greedy to stop exploring, or New network to start from scratch. Keys: T or Space for speed, G for greedy, R for a new network.

  • Deep Q-learning
  • Experience replay
  • Double DQN target network
  • Backpropagation with Adam
  • Flood-fill state features
  • Dot-matrix LCD rendering

View the source · one module, plus a small shared runtime for sizing, the animation loop and input

Build your own

Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.

Build a snake game that a deep Q-network learns to play by itself, with JavaScript and the HTML canvas element. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.

Start simple:
- Make a 12x12 snake game: a snake three cells long, food in a random free cell, death on hitting a wall or itself.
- Give the agent three actions relative to its heading (turn left, go ahead, turn right) and a small state vector: for each action, whether that square is deadly, plus four flags for whether the food is ahead, behind, left or right of the head.
- Write a tiny neural network by hand: inputs, one hidden layer of 32 ReLU units, three linear outputs (one Q-value per action). Train it with plain gradient descent on the squared error.
- Use Q-learning with experience replay: play with an epsilon-greedy policy (epsilon decays from 1 to 0.01 over a few thousand steps), store every (state, action, reward, next state, done) in a ring buffer, and after each step train on a random minibatch of 32 toward reward + 0.9 * max Q(next state). Rewards: +1 for food, -1 for dying.
- Run many training steps per frame on a hidden game, show a second game played by the current network at a watchable speed, and chart the score per episode.

Once that works, make it beautiful:
- Add a target network (a copy of the weights refreshed every few hundred updates) and the Adam optimizer for steadier learning.
- Draw the board like a reflective green dot-matrix LCD, with faint unlit pixels and a soft drop shadow under lit ones.
- Show the three Q-values as arrows around the snake's head, sized by preference.

Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as flood-fill features so it stops trapping itself, letting me take the controls with the arrow keys, or visualizing the network's activations live.
PreviousSector EngineA 2.5D portal engine in the mid-90s style: a slow walk round a castle at golden hour. NextBorsuk-Ulam WeatherA topology theorem guarantees two antipodes share the weather. Watch them get found.

Related visualizations

  • Neural NetworkMachine learning A tiny neural network learns to classify points, live, in your browser.
  • Perfect SnakeAlgorithms A snake AI that can never crash, cutting safe shortcuts until it fills every cell.
  • Topological SnakeGames Snake on a torus, a Klein bottle and a projective plane, with the real surface turning beside it.
  • Growing Neural CellsMachine learning Cells running one tiny learned rule grow a gecko from a single seed and heal its wounds.
  • Lottery TicketMachine learning Prune a network to a sparse winning ticket that still learns, round by round.
  • ReLU Stained GlassMachine learning A ReLU network's exact linear regions, drawn as a stained glass window.
  • CPPN GardenMachine learning Breed images grown by tiny neural networks by picking the ones you like.
  • Self-Organizing QuiltMachine learning Kohonen maps sew a quilt from fabric dyes and drape a sheet over 3D point clouds.
  • GrokkingMachine learning A tiny network memorizes modular addition, then suddenly understands it.

Use ← and → to move between demos. While the canvas has focus, keys go to the demo instead.

← More from Emergent Mind Labs