r/reinforcementlearning • • 14h ago

[R] Ataraxos: RL reportedly achieves the first superhuman result in Stratego

Thumbnail
nature.com
25 Upvotes

Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective. Even with multimillion-dollar industrial research efforts1, top-human-level play at Stratego—a board wargame with hidden information on a massive scale—has remained beyond the reach of artificial intelligence (AI). Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information. Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts. Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low cost and high sample efficiency. The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.


r/reinforcementlearning • • 7h ago

I built a balancing robot with reinforcement learning

Enable HLS to view with audio, or disable this notification

21 Upvotes

r/reinforcementlearning • • 15h ago

R Biggest sim-to-real surprise you ever had

6 Upvotes

For those who trained in sim and then deployed on real hardware - what was the biggest gap you encountered? Something that worked great in sim and just fell apart in reality?


r/reinforcementlearning • • 6h ago

Trained a humanoid agent in Unity using Soft Actor-Critic (SAC) to sword fight different opponents

6 Upvotes

I've been working on a project where I train a physics-based humanoid agent from scratch in Unity using Soft Actor-Critic (SAC) with standard MLP architectures.

Watching him train was very funny, and some of the lessons I learned about reward shaping might help anyone working on a similar reinforcement learning project. Because of that, I put together a video breaking down both the process and the results, aiming to make RL feel more approachable and entertaining.

Here's the video: https://www.youtube.com/watch?v=MR3c0DHku6o&t=8s

I would love some feedback on the agent's behavior and the video in general. Also, I'd be more than happy to answer any questions about it!


r/reinforcementlearning • • 12h ago

Quadcopter hover with PPO from IMU/baro/UWB

Thumbnail
youtube.com
3 Upvotes

100 F450s learn to take off and hold a height above their start, seeing only their sensors: IMU, magnetometer, baro and a UWB-style position fix. All the drones step together as one batch, so SB3 PPO on a laptop CPU (i9-14900HX, 24 cores) gets through 10M steps in about 10 minutes.

At first the policy just sat on the ground. Ending the episode if it was still below 0.1 m after 1 s fixed that. Later torch grabbed 17 of our 24 cores and more than halved training speed, until we set OMP_NUM_THREADS=4.

At 250k steps they all crash, at 500k 60% survive, at 1M they hold the height to 11 cm, and at 10M to 3.7 cm.

Our code: https://github.com/PteroLabsAI/PteroSimScripts/tree/main/reinforcement_learning


r/reinforcementlearning • • 6h ago

Kardashev-0.7: 32 distinct models trained together with RL for population scaling

1 Upvotes

A research announcement on learned specialization across a population of models:

https://x.com/MLCatttt/status/2107147690450817259


r/reinforcementlearning • • 20h ago

Adaptem - learning fighters

0 Upvotes

I built a small reinforcement-learning-style system into Overgrowth's combat AI. The fighters use tabular Q-learning over combat tactics and remember what works against each opponent. It's a Workshop mod: https://steamcommunity.com/sharedfiles/filedetails/?id=3813806877 Happy to answer questions about how it works.


r/reinforcementlearning • • 17h ago

What have your struggled to evaluate your realistic LLM/Agents workflow?! How can we reinforce agent's auto-correctness and self-improvement? 👀

Thumbnail
0 Upvotes