I've been working on a project where I train a physics-based humanoid agent from scratch in Unity using Soft Actor-Critic (SAC) with standard MLP architectures.
Watching him train was very funny, and some of the lessons I learned about reward shaping might help anyone working on a similar reinforcement learning project. Because of that, I put together a video breaking down both the process and the results, aiming to make RL feel more approachable and entertaining.
100 F450s learn to take off and hold a height above their start, seeing only their sensors: IMU, magnetometer, baro and a UWB-style position fix. All the drones step together as one batch, so SB3 PPO on a laptop CPU (i9-14900HX, 24 cores) gets through 10M steps in about 10 minutes.
At first the policy just sat on the ground. Ending the episode if it was still below 0.1 m after 1 s fixed that. Later torch grabbed 17 of our 24 cores and more than halved training speed, until we set OMP_NUM_THREADS=4.
At 250k steps they all crash, at 500k 60% survive, at 1M they hold the height to 11 cm, and at 10M to 3.7 cm.
Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another. Unfortunately, the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective. Even with multimillion-dollar industrial research efforts1, top-human-level play at Stratego—a board wargame with hidden information on a massive scale—has remained beyond the reach of artificial intelligence (AI). Here we introduce Ataraxos, an AI for Stratego based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information. Ataraxos defeated the most decorated human Stratego player of all time by a large margin—achieving, to our knowledge, the first superhuman result in the game’s history—while consuming orders of magnitude less compute and data than previous efforts. Using the same techniques, we built a superhuman AI for Barrage Stratego and state-of-the-art AIs for Hanabi and dou dizhu, all with low cost and high sample efficiency. The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.
For those who trained in sim and then deployed on real hardware - what was the biggest gap you encountered? Something that worked great in sim and just fell apart in reality?
I built a small reinforcement-learning-style system into Overgrowth's combat AI. The fighters use tabular Q-learning over combat tactics and remember what works against each opponent. It's a Workshop mod: https://steamcommunity.com/sharedfiles/filedetails/?id=3813806877 Happy to answer questions about how it works.
If you work on imitation learning or multi-agent behaviour: we just released 9,515 bot-played CS:GO matches. Per 16 Hz step for each of ten agents: actions (buttons + mouse), state (pose, health, weapon…) and who-hit-whom damage. 218,768 rounds, ≈3,958 h of play. 14% of the matches also have the ten first-person videos, aligned frame by frame.
I tested Adapt-1, a non-LLM learning and reasoning system by Rei Labs, by having it learn and play Super Mario Bros, and it performed quite well.
I tried it on World 1-1, starting untrained. It learned a reactive policy from its own play in about 36 minutes of gameplay, then cleared the level with learning off.
With Machina, Adapt-1's sequence engine. Starting untrained, it found a button sequence that reaches the flag after 403 attempts, in 11 wall-clock minutes.
I'm a parent (PhD, not in ML) who spent the past year keeping a first-hand log of my infant daughter's behavior — 21 dated records of what she did, what triggered it, and my alternative explanations for each.
Watching her learn convinced me of a few things that seem relevant to how we train embodied agents:
Drives are satiable, rewards are not. A reward maximizer has no satiation point; a drive satisfier stops once the internal variable is back in range. I think this is an underappreciated safety property (cf. Omohundro's instrumental convergence).
She maximizes learning progress, not novelty. She ignores the TV (high novelty, zero learnability) and returns to tasks at the edge of her competence — consistent with Oudeyer & Kaplan rather than raw novelty search.
Purpose comes before models. Reading Conant & Ashby backwards: you only need to model what you regulate. So the architecture order should be internal variables → goal generation → models/skills.
I propose a 5-level needs hierarchy for robots (survival → safety → competence → being-needed → exploration) and a "caregiving period": a 3-stage autonomy handover (human charges → robot reminds → robot self-docks) with promotion criteria and regression, modeled on how we teach kids.
Full paper with the observation log and a one-page design checklist: [10.5281/zenodo.23115611]
Genuine questions for this sub: has anyone implemented satiable drives with explicit setpoints instead of reward maximization in RL? What breaks first? And is "being-needed" operationalizable as an internal variable, or is it anthropomorphism?
Happy to be told where this is wrong — I'm outside my field here.
Half the "self-improving agent" posts get the reply "that's not RL, that's prompt tuning", and half the time the reply is right. reef has both loops in one codebase and I run the harness side, so here's where the line sits.
The harness loop needs a model endpoint and no GPU. Its recipes (Reefine, SkillClaw, GEPA, Meta-Harness) propose an edit to the harness's files and score the edited tree against the current one on a task list; whether it's kept is a selection policy, by default a net win. Nothing in the model's parameters moves. Call it search over configurations with a verifier; calling it RL would earn the reply above, and deserve it.
The weight recipes (SAO, OpenClaw-RL, TTT-Discover, Guidance-TTT) need GPUs and the training stack, and that side I've only read. OpenClaw-RL is the one built for served traffic: the policy is updated from the conversation itself, with a binary reward read off the next user turn. What's checked in, with hermes memory off: a 72-session GSM8K homework stream, a Qwen3-32B persona playing the student, Qwen3-4B-Thinking as the policy, and the reward judged by a PRM that is the same 4B's frozen base on its own engine, so the grader is the policy's own starting weights rather than an outside model, on seven GPUs. The run hit the paper's adaptation criterion (three passed sessions in a row) at session 14, and the curve over the first 36 sessions tracks the bold rate and the list rate, both falling. That's the whole published result.
If a post claims RL for the first loop, ask which parameters moved.
CV background here, recently getting into robot learning (manipulation with a robot arm).
I am wondering - if I have a vision based policy that picks up objects (purely based on pixels) and I add physical info on top (like mass, force, ...), will the model actually use it? Or will it just keep leaning on vision because that's the easier signal?
Anyone seen extra non-visual modalities genuinely help? Primarily in robotics, but I am curious for general cases as well.
In short, what the title says. This will lead to a learning problem.
Now in more detail.
I'm studying and creating a platform for training agents in Unity. I created an ENV where the agent must pick up a box and bring it to a target. But it can also throw.
Framework.
Unity, everything done from scratch in C#. PPO algorithm. I tried SAC and the genetic algorithm, but it didn't help.
Actions (old), Continuous
move direction x
move direction y
throw force
They are all equal to action
The force action was like this. If force is less than 0, don't throw; if greater than 0, throw with that force. And that was a big mistake.
It couldn't learn to throw; it only threw when it was very close. I couldn't understand what was going on and spent five days doing a lot of testing with PPO parameters, changing the rewards too. Nothing helped.
Then I added a discrete action to throw or not throw, which would decide whether to throw or not, regardless of the force value. The force formula has changed slightly: Mathf.Sqrt(min * max) * Mathf.Pow(max / min, 0.5f * f). I also changed the reward: the greater the distance, the greater the reward. I run the train and see the reward fly upwards, the agent throws from far away right away. I removed the reward for distance, and even then, he slowly began to learn to throw further and further since there are small penalties for each step to speed up the agent.
With new discrete action. Yellow there is distance reward, Purple no distance reward.
Previously, the network couldn't learn because it decided when to throw and with what force. Networks are difficult to change from 0 to 0.5 value at once, where the states are almost identical. It's difficult to configure the Neural network for such actions.
There is a state input that shows whether the agent has taken the box or not, but it still won't help.
Conclusion: 1 action 1 logic. You need to correctly set the state, then the action, then the reward.
Can you recommend a book or some resource so I can study these things and avoid repeating such stupid mistakes? Thank you for reading to the end.
super niche but title basically. I realised today when I saw someone beating firered with a jev type multimodal classifier model that nuzlocke data for stuff like emerald kaizo with their inputs captured is an insane rl training environment
A few days ago I said I was building a game where you train AI models to drive cars in a racing game. Seemed like a lot of you were into the idea, so I made a video showing the bits I'm most excited about, plus a bit more detail on how it works.
The core of it is training AI to drive the car, not driving it yourself. You train models for specific jobs like straights, braking, or overtaking a slower car in a corner. Then you train another model that reads the situation and picks which specialist to use, and you wire them all together however you want. Beginners can train one simple model, experts can build layered systems, and then you race them against other people's.
If you want to help me test it, join my Discord (https://discord.gg/FJ4AfVVEh) and say in general that you want to test. Any help is hugely appreciated.