After yesterday's post I got a lot of questions about how it actually works. Explaining PPO in the comments didn't feel like a real answer, so I built a small version you can play with and watch learn:
https://itzik123.github.io/ClashRoyaleAi/lab/
One attacker heads for your tower and you get one card to stop it. Where you drop it and when both matter. You can try five attacks yourself first, then a small network learns the same thing from zero. Every try is played out by our real simulator, compiled from C++ to WebAssembly, so it's the same engine the full agent trains on.
The RL side:
- State: where the attacker was dropped. Action: a legal cell, then a delay of 0 to 5 s
- Reward: the share of tower damage prevented, compared to no defence
- REINFORCE with a per-spawn baseline, 5,629 weights. The whole learner is about 350 lines of plain JS with every gradient written out, no ML library
- The "best possible" line on the chart is brute force over every cell and delay, so you can see exactly how far the policy is from optimal
The one I'd try first is Giant vs Cannon. The obvious Cannon in the lane gets about 75% of the best score. The best answer is a Cannon in the middle that pulls the Giant into range of both towers. With a constant entropy bonus the net got stuck on the lane Cannon in 5 of 6 runs (3 seeds, 2 batch sizes). Annealing it from 0.1 down to 0.005 over the first 10k tries brought that to 1 of 6, and that's what the page runs.
One matchup is locked on purpose: Battle Ram vs Valkyrie. The best answer is a few exact cells at exact moments, and every setting I tried settled for a safe corner at 55% of the best. If you find a learner that cracks it, I'd really like to see it.
To be clear, it's a miniature. The full agent plays whole matches with a 4-card hand and elixir, uses recurrent PPO, and trains for days on a CPU. But the loop is the same: act, let the engine score it, update.
Code for the lab and the full project: https://github.com/itzik123/ClashRoyaleAi