A deep Q-network learns snake from scratch, live, on a retro dot-matrix LCD.
A small neural network (17 inputs, two hidden layers, 3 outputs) learns to play snake by deep Q-learning, starting from random weights every time the page loads. It sees the board from the snake's point of view: for each of its three moves, whether that square is deadly, how far it can see, and how much of the board a flood fill can still reach from there, plus where the food is. Every move goes into a 40,000-slot replay memory, and each step a random minibatch of 32 is replayed to pull Q(s, a) toward the reward plus the discounted value of the next state, using Double DQN targets from a frozen target network that is refreshed every 250 updates. A hidden training snake takes as many steps as fit in a few milliseconds each frame, so within seconds the average score climbs from zero into the twenties or thirties, while the snake on the LCD plays with the same network and its arrows show the Q-value of turning left, going ahead or turning right.
Try it. Take over with the arrow keys, WASD or a swipe (your moves are learned from too, and the agent's arrows still show what it would do); it hands back to the agent when you stop. Tap Training to switch between Slow, Fast and Turbo, Greedy to stop exploring, or New network to start from scratch. Keys: T or Space for speed, G for greedy, R for a new network.
Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.
Build a snake game that a deep Q-network learns to play by itself, with JavaScript and the HTML canvas element. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.
Start simple:
- Make a 12x12 snake game: a snake three cells long, food in a random free cell, death on hitting a wall or itself.
- Give the agent three actions relative to its heading (turn left, go ahead, turn right) and a small state vector: for each action, whether that square is deadly, plus four flags for whether the food is ahead, behind, left or right of the head.
- Write a tiny neural network by hand: inputs, one hidden layer of 32 ReLU units, three linear outputs (one Q-value per action). Train it with plain gradient descent on the squared error.
- Use Q-learning with experience replay: play with an epsilon-greedy policy (epsilon decays from 1 to 0.01 over a few thousand steps), store every (state, action, reward, next state, done) in a ring buffer, and after each step train on a random minibatch of 32 toward reward + 0.9 * max Q(next state). Rewards: +1 for food, -1 for dying.
- Run many training steps per frame on a hidden game, show a second game played by the current network at a watchable speed, and chart the score per episode.
Once that works, make it beautiful:
- Add a target network (a copy of the weights refreshed every few hundred updates) and the Adam optimizer for steadier learning.
- Draw the board like a reflective green dot-matrix LCD, with faint unlit pixels and a soft drop shadow under lit ones.
- Show the three Q-values as arrows around the snake's head, sized by preference.
Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as flood-fill features so it stops trapping itself, letting me take the controls with the arrow keys, or visualizing the network's activations live.