Epsilon-greedy, UCB1 and Thompson sampling race to find the best slot machine.
Five pastel slot machines pay out with hidden probabilities, and three players work the same floor at once. Epsilon-greedy plays the best average except for 10% random pulls, UCB1 plays the best average plus an optimism bonus of sqrt(2 ln t / n), and Thompson sampling draws a plausible payout rate from each machine's Beta posterior (via two Marsaglia-Tsang gamma draws) and plays the best draw. Above every machine, each player's Beta(1 + wins, 1 + losses) posterior is drawn as a live curve with its decision value marked underneath, and the race at the bottom tracks cumulative pseudo-regret. Partway through the night the house quietly turns the worst machine into the best one, so you can see which strategy notices; left alone, the players then switch to a fading memory of discounted counts and re-adapt.
Try it. Click or tap a machine (or press 1 to 5) to secretly change its payout rate. Press D to give the players a fading memory with discounted counts, which helps them re-adapt, plus and minus to change the speed, and R to start a new night with new machines.
Paste this into Claude Code, Codex or any coding agent to get a simple version running, then take it wherever you like.
Build a multi-armed bandit simulation with JavaScript and the HTML canvas element, comparing epsilon-greedy, UCB1 and Thompson sampling. Put everything in a single index.html file with no libraries or build step, so I can open it directly in a browser.
Start simple:
- Make a canvas that fills the window, stays sharp on high-DPI screens (scale by devicePixelRatio), and resizes with the window.
- Create five slot machines with hidden payout probabilities between 0.2 and 0.7. A pull pays 1 with that probability, else 0.
- Write three players that each keep wins and losses per machine: epsilon-greedy (play the best average, but 10% of the time a random machine), UCB1 (play every machine once, then the best average plus sqrt(2 ln t / n)), and Thompson sampling (draw a rate from Beta(1 + wins, 1 + losses) for each machine and play the largest draw).
- To sample a Beta, sample two gamma variables with the Marsaglia and Tsang method and return x / (x + y).
- Each step, every player pulls once. Track cumulative regret: the best machine's rate minus the rate of the machine played, summed over time, and plot the three regret curves.
Once that works, make it beautiful:
- Draw the machines as friendly pastel slot machines in a row, with a colored chip for each player sitting at the machine it just played.
- Above each machine, plot every player's Beta posterior density for that machine as a filled curve (use a log-gamma function to normalize it), with a dashed line at the true rate.
- Start slow so single pulls are visible, then speed up. Let me click a machine to secretly change its payout rate and watch who notices.
Explain the key ideas in short code comments. When you're done, tell me how to open it and suggest three directions I could take it next, such as discounted counts for changing payouts, a contextual bandit, or comparing many random casinos to average out luck.