Fifty ghost landers learn soft touchdowns by natural evolution strategies.
One small neural network (8 inputs, 12 tanh units, throttle and torque out) flies every lander, and no gradient is ever backpropagated. Each generation perturbs its 122 weights along 25 random directions, both ways (antithetic sampling), flies all 50 variants through the same three approaches, turns their returns into centered ranks, and lets Adam step along the rank-weighted sum of directions: natural evolution strategies, as used to train robots at scale. Training runs headlessly, dozens of generations a second. Every generation's 50 flight paths are added into an offscreen canvas with additive blending and slowly fade, so the screen becomes a long exposure of wild early arcs dimming beneath bright filaments converging on the pad, while the latest ghosts and the mean policy (solid) fly live in front.
Try it. Drag the gravity and wind sliders (or use the arrow keys) and watch the photograph record the policy being knocked off course and recovering. Click the sky to launch the ghosts from there, click the ground to move the pad. W tours the worlds, F speeds up the replay, R starts learning from scratch.
Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.
Build a lunar lander that teaches itself to land using evolution strategies, with JavaScript and the HTML canvas element. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.
Start simple:
- Simulate a 2D lander (position, velocity, angle, spin) under lunar gravity, 1.62 m/s², at a fixed 30 steps per second. It has a main engine (throttle 0 to 1, about 2.2 times gravity at full power, pushing along its up direction) and a turning torque.
- Draw flat ground with a landing pad and a simple lander shape with a flame that scales with the throttle.
- The pilot is a tiny neural network: inputs are distance to the pad, height, both velocities, angle, spin and a bias of 1; 12 tanh hidden units; two outputs (sigmoid for throttle, tanh for torque). Store all weights in one flat array.
- Score an episode when it touches down: penalize distance from the pad, speed, tilt and spin, and add a bonus for a soft landing on the pad.
- Train with evolution strategies: each generation, draw 25 random noise vectors, evaluate the weights plus and minus 0.1 times each one on the same starting positions, rank the scores, and move the weights a small step along the rank-weighted sum of the noise vectors. Run several generations per frame without drawing them.
- Each frame, fly the current weights live and show the generation and landing rate.
Once that works, make it beautiful:
- Record every flight path and draw them all into an offscreen canvas with additive blending and a slow fade, so the screen becomes a long-exposure photograph of trajectories converging on the pad.
- Paint a starry sky, cratered terrain and blinking pad lights, and shift the trail color as the generations pass.
- Add sliders for gravity and wind so you can watch the policy get knocked off course and adapt.
Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as using Adam for the update step, comparing against a plain genetic algorithm, or giving the lander fuel limits and uneven terrain.