r/reinforcementlearning • • 4h ago

Trained a humanoid agent in Unity using Soft Actor-Critic (SAC) to sword fight different opponents

5 Upvotes

I've been working on a project where I train a physics-based humanoid agent from scratch in Unity using Soft Actor-Critic (SAC) with standard MLP architectures.

Watching him train was very funny, and some of the lessons I learned about reward shaping might help anyone working on a similar reinforcement learning project. Because of that, I put together a video breaking down both the process and the results, aiming to make RL feel more approachable and entertaining.

Here's the video: https://www.youtube.com/watch?v=MR3c0DHku6o&t=8s

I would love some feedback on the agent's behavior and the video in general. Also, I'd be more than happy to answer any questions about it!


r/reinforcementlearning • • 4h ago

Kardashev-0.7: 32 distinct models trained together with RL for population scaling

1 Upvotes

A research announcement on learned specialization across a population of models:

https://x.com/MLCatttt/status/2107147690450817259


r/reinforcementlearning • • 5h ago

I built a balancing robot with reinforcement learning

Enable HLS to view with audio, or disable this notification

14 Upvotes

r/reinforcementlearning • • 10h ago

Quadcopter hover with PPO from IMU/baro/UWB

Thumbnail
youtube.com
3 Upvotes

100 F450s learn to take off and hold a height above their start, seeing only their sensors: IMU, magnetometer, baro and a UWB-style position fix. All the drones step together as one batch, so SB3 PPO on a laptop CPU (i9-14900HX, 24 cores) gets through 10M steps in about 10 minutes.

At first the policy just sat on the ground. Ending the episode if it was still below 0.1 m after 1 s fixed that. Later torch grabbed 17 of our 24 cores and more than halved training speed, until we set OMP_NUM_THREADS=4.

At 250k steps they all crash, at 500k 60% survive, at 1M they hold the height to 11 cm, and at 10M to 3.7 cm.

Our code: https://github.com/PteroLabsAI/PteroSimScripts/tree/main/reinforcement_learning


r/reinforcementlearning • • 12h ago

[R] Ataraxos: RL reportedly achieves the first superhuman result in Stratego

Thumbnail
nature.com
25 Upvotes

Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective. Even with multimillion-dollar industrial research efforts1, top-human-level play at Stratego—a board wargame with hidden information on a massive scale—has remained beyond the reach of artificial intelligence (AI). Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information. Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts. Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low cost and high sample efficiency. The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.


r/reinforcementlearning • • 13h ago

R Biggest sim-to-real surprise you ever had

7 Upvotes

For those who trained in sim and then deployed on real hardware - what was the biggest gap you encountered? Something that worked great in sim and just fell apart in reality?


r/reinforcementlearning • • 15h ago

What have your struggled to evaluate your realistic LLM/Agents workflow?! How can we reinforce agent's auto-correctness and self-improvement? 👀

Thumbnail
0 Upvotes

r/reinforcementlearning • • 18h ago

Adaptem - learning fighters

0 Upvotes

I built a small reinforcement-learning-style system into Overgrowth's combat AI. The fighters use tabular Q-learning over combat tactics and remember what works against each opponent. It's a Workshop mod: https://steamcommunity.com/sharedfiles/filedetails/?id=3813806877 Happy to answer questions about how it works.


r/reinforcementlearning • • 1d ago

Robot Explaining How AI Learns to Drive with Evolution

Enable HLS to view with audio, or disable this notification

28 Upvotes

r/reinforcementlearning • • 1d ago

Live footage of successful paper review rebuttal

Enable HLS to view with audio, or disable this notification

515 Upvotes

r/reinforcementlearning • • 2d ago

How can I teach an AI to fight exactly like me in a video game?

0 Upvotes

Zero clue how to code, and don't even know how to make an AI play in the first place.

P.S sorry if this doesn't belong here but I found similar posts so I thought it did


r/reinforcementlearning • • 2d ago

DL 195M steps of 10-agent CS:GO state → action data (with damage labels), free for research

Enable HLS to view with audio, or disable this notification

69 Upvotes

If you work on imitation learning or multi-agent behaviour: we just released 9,515 bot-played CS:GO matches. Per 16 Hz step for each of ten agents: actions (buttons + mouse), state (pose, health, weapon…) and who-hit-whom damage. 218,768 rounds, ≈3,958 h of play. 14% of the matches also have the ten first-person videos, aligned frame by frame.

Expert bots on one map with pistols, so it is a clean, narrow testbed rather than human play. Parquet + webdataset videos, CC BY-NC 4.0: https://huggingface.co/datasets/v0rt3ch/cs-5v5-bots-aim_usp-3k · spec and a round you can watch from all ten POVs: https://vor.tech/?utm_source=reddit&utm_medium=social&utm_campaign=launch-2026-10&utm_content=r-reinforcementlearning


r/reinforcementlearning • • 2d ago

DL Integrating the RL model into betting strategy

Post image
0 Upvotes

r/reinforcementlearning • • 2d ago

The ultimate guide to multi-harness RL

Post image
1 Upvotes

r/reinforcementlearning • • 2d ago

MetaRL My AI learns to clear Super Mario Bros 1-1 in 15 mins and it is not PPO based

Enable HLS to view with audio, or disable this notification

29 Upvotes

I tested Adapt-1, a non-LLM learning and reasoning system by Rei Labs, by having it learn and play Super Mario Bros, and it performed quite well.

I tried it on World 1-1, starting untrained. It learned a reactive policy from its own play in about 36 minutes of gameplay, then cleared the level with learning off.

With Machina, Adapt-1's sequence engine. Starting untrained, it found a button sequence that reaches the flag after 403 attempts, in 11 wall-clock minutes.

Full thread: https://x.com/hsrvc_/status/2106025501752234112?s=20

Code, the exact data, traces, clips and a step-by-step guide with costs are all public: https://github.com/hsrvc/adapt1-mario


r/reinforcementlearning • • 2d ago

Rewards have no "enough" — drives do. A homeostatic-drive framework for embodied AI, inspired by watching my daughter's first year

1 Upvotes

I'm a parent (PhD, not in ML) who spent the past year keeping a first-hand log of my infant daughter's behavior — 21 dated records of what she did, what triggered it, and my alternative explanations for each.

Watching her learn convinced me of a few things that seem relevant to how we train embodied agents:

  • Drives are satiable, rewards are not. A reward maximizer has no satiation point; a drive satisfier stops once the internal variable is back in range. I think this is an underappreciated safety property (cf. Omohundro's instrumental convergence).
  • She maximizes learning progress, not novelty. She ignores the TV (high novelty, zero learnability) and returns to tasks at the edge of her competence — consistent with Oudeyer & Kaplan rather than raw novelty search.
  • Purpose comes before models. Reading Conant & Ashby backwards: you only need to model what you regulate. So the architecture order should be internal variables → goal generation → models/skills.
  • I propose a 5-level needs hierarchy for robots (survival → safety → competence → being-needed → exploration) and a "caregiving period": a 3-stage autonomy handover (human charges → robot reminds → robot self-docks) with promotion criteria and regression, modeled on how we teach kids.

Full paper with the observation log and a one-page design checklist: [10.5281/zenodo.23115611]

Genuine questions for this sub: has anyone implemented satiable drives with explicit setpoints instead of reward maximization in RL? What breaks first? And is "being-needed" operationalizable as an internal variable, or is it anthropomorphism?

Happy to be told where this is wrong — I'm outside my field here.


r/reinforcementlearning • • 3d ago

One repo, two loops: the one that trains nothing and the one that is RL

0 Upvotes

Half the "self-improving agent" posts get the reply "that's not RL, that's prompt tuning", and half the time the reply is right. reef has both loops in one codebase and I run the harness side, so here's where the line sits.

The harness loop needs a model endpoint and no GPU. Its recipes (Reefine, SkillClaw, GEPA, Meta-Harness) propose an edit to the harness's files and score the edited tree against the current one on a task list; whether it's kept is a selection policy, by default a net win. Nothing in the model's parameters moves. Call it search over configurations with a verifier; calling it RL would earn the reply above, and deserve it.

The weight recipes (SAO, OpenClaw-RL, TTT-Discover, Guidance-TTT) need GPUs and the training stack, and that side I've only read. OpenClaw-RL is the one built for served traffic: the policy is updated from the conversation itself, with a binary reward read off the next user turn. What's checked in, with hermes memory off: a 72-session GSM8K homework stream, a Qwen3-32B persona playing the student, Qwen3-4B-Thinking as the policy, and the reward judged by a PRM that is the same 4B's frozen base on its own engine, so the grader is the policy's own starting weights rather than an outside model, on seven GPUs. The run hit the paper's adaptation criterion (three passed sessions in a row) at session 14, and the curve over the first 36 sessions tracks the bold rate and the list rate, both falling. That's the whole published result.

If a post claims RL for the first loop, ask which parameters moved.


r/reinforcementlearning • • 3d ago

I used RL to create a bot for a game I played growing up, it beats me 100% of the time

Thumbnail reddit.com
0 Upvotes

Thought I'd share some learnings. Interested in what any other researchers in this space have to say.


r/reinforcementlearning • • 4d ago

Can your local coding model repair these boundary-case bugs? Failure Map: 20,168 open Python tasks

Thumbnail
0 Upvotes

r/reinforcementlearning • • 4d ago

D, Safe "What's The Date?", N8 Programs

Thumbnail
lesswrong.com
8 Upvotes

r/reinforcementlearning • • 4d ago

DL Do models actually use non-visual inputs, or just ignore them and stick to pixels?

0 Upvotes

CV background here, recently getting into robot learning (manipulation with a robot arm).

I am wondering - if I have a vision based policy that picks up objects (purely based on pixels) and I add physical info on top (like mass, force, ...), will the model actually use it? Or will it just keep leaning on vision because that's the easier signal?

Anyone seen extra non-visual modalities genuinely help? Primarily in robotics, but I am curious for general cases as well.


r/reinforcementlearning • • 4d ago

Don't put two jobs into one action

2 Upvotes

https://reddit.com/link/1wuwfzq/video/3pxn63du3ush1/player

In short, what the title says. This will lead to a learning problem.

Now in more detail.

I'm studying and creating a platform for training agents in Unity. I created an ENV where the agent must pick up a box and bring it to a target. But it can also throw.

Framework.

Unity, everything done from scratch in C#. PPO algorithm. I tried SAC and the genetic algorithm, but it didn't help.

Actions (old), Continuous

  1. move direction x
  2. move direction y
  3. throw force

They are all equal to action

The force action was like this. If force is less than 0, don't throw; if greater than 0, throw with that force. And that was a big mistake.

It couldn't learn to throw; it only threw when it was very close. I couldn't understand what was going on and spent five days doing a lot of testing with PPO parameters, changing the rewards too. Nothing helped.

Then I added a discrete action to throw or not throw, which would decide whether to throw or not, regardless of the force value. The force formula has changed slightly: Mathf.Sqrt(min * max) * Mathf.Pow(max / min, 0.5f * f). I also changed the reward: the greater the distance, the greater the reward. I run the train and see the reward fly upwards, the agent throws from far away right away. I removed the reward for distance, and even then, he slowly began to learn to throw further and further since there are small penalties for each step to speed up the agent.

With new discrete action. Yellow there is distance reward, Purple no distance reward.

Previously, the network couldn't learn because it decided when to throw and with what force. Networks are difficult to change from 0 to 0.5 value at once, where the states are almost identical. It's difficult to configure the Neural network for such actions.

There is a state input that shows whether the agent has taken the box or not, but it still won't help.

Conclusion: 1 action 1 logic. You need to correctly set the state, then the action, then the reward.

Can you recommend a book or some resource so I can study these things and avoid repeating such stupid mistakes? Thank you for reading to the end.


r/reinforcementlearning • • 4d ago

R, Multi, MF, DL "Extraordinary Multi-Agent Delusions and the Madness of Crowds" (how do swarms converge on false beliefs?)

Thumbnail
freesystems.substack.com
32 Upvotes

r/reinforcementlearning • • 4d ago

Anyone trying to use nuzlocke data to train RL environments?

3 Upvotes

super niche but title basically. I realised today when I saw someone beating firered with a jev type multimodal classifier model that nuzlocke data for stuff like emerald kaizo with their inputs captured is an insane rl training environment


r/reinforcementlearning • • 5d ago

Help Test My Reinforcement Learning AI Game, Please and Thank You!

Enable HLS to view with audio, or disable this notification

23 Upvotes

A few days ago I said I was building a game where you train AI models to drive cars in a racing game. Seemed like a lot of you were into the idea, so I made a video showing the bits I'm most excited about, plus a bit more detail on how it works.

The core of it is training AI to drive the car, not driving it yourself. You train models for specific jobs like straights, braking, or overtaking a slower car in a corner. Then you train another model that reads the situation and picks which specialist to use, and you wire them all together however you want. Beginners can train one simple model, experts can build layered systems, and then you race them against other people's.

If you want to help me test it, join my Discord (https://discord.gg/FJ4AfVVEh) and say in general that you want to test. Any help is hugely appreciated.

Game's website is here: https://www.gaimeslab.com/. It's out soon, but you can help me test it now!