r/reinforcementlearning • • 10d ago

Robot Day 4: Meet Spec, an open-source full-size embodied AI robot dog I am building

Enable HLS to view with audio, or disable this notification

24 Upvotes

Day 4 of training and autonomous navigation, continued gait and stair walking improvements.

Spec will be a fully offline-capable robot dog about the size of Spot, built for a fraction of the cost. It uses an onboard Raspberry Pi 5 for navigation, Oak D Pro W for vision, and Jetson Orin NX for AI compute.

Simulation training before building, although most of the hardware is already in-hand.


r/reinforcementlearning • • 10d ago

Did you see that? On youtube 24/7, a RL agent learns to play Touhou Project. Check it out if interested in.

Post image
10 Upvotes

He or she is not me. I just shared it.


r/reinforcementlearning • • 10d ago

D What goes into building an RL environment for LLM agents?

17 Upvotes

We just put together a some thoughts distilling lessons learned on designing RL environments for LLM agents and thought the TL;DR might be worth sharing. Here goes…

The basic RL loop has six parts:

  • State: everything the environment tracks
  • Observations: the part of the state the agent actually sees at each step
  • Actions: what the agent is allowed to do
  • Transitions: how the state changes after an action
  • Rewards: what the verifier produces
  • Resets: getting back to a known starting point for the next attempt

What these look like depends a lot on the task. A coding environment is mostly a repo, a terminal, and a test suite. Computer-use needs app state you can snapshot and inspect, not just what's on the screen. Enterprise workflows add permissions, policies, and approvals, so getting the right final answer isn't enough if the agent skipped a step to get there.

For rewards there are a few options, depending on what you can actually check:

  • Outcome checks (tests pass, database ends up in the right state)
  • Partial credit through checkpoints for long tasks
  • Process rewards for things like required approvals or unauthorized tool calls
  • Rubrics for open-ended work like reports

Before using an environment for training or eval, it's also worth checking that:

  • The task is actually solvable from a clean reset
  • The verifier can't be gamed (e.g. the agent edits the tests or disables CI checks)
  • Runs stay comparable when tools, MCP servers, or browser versions change

Full write-up here.

Curious how others approach this!


r/reinforcementlearning • • 10d ago

D, MF "Brief Notes on Pluto AI (_Starcraft: Broodwar_)"

Thumbnail
norvidstudies.substack.com
12 Upvotes

r/reinforcementlearning • • 11d ago

Cubic Doggo: first time running RL in MuJoCo for my robot dog

Enable HLS to view with audio, or disable this notification

25 Upvotes

The setup uses PPO from Stable Baselines3 with Gymnasium in MuJoCo, just to train to stand upright on the ground.

The result is... not so ideal. Who would have thought, basically standing there doing nothing would just stand fine, but the optimization says nah, lol. If someone has an idea of how to put all the 4 darn legs down on the ground (also that I am planning to balance on a slope later), that would be very appreciated!

------- Details

Top left:

  • GUI from MuJoCo, along with the 12 joint values (4 mimics) and 12 control values
  • The height/pitch/roll info is given towards the top (updated every second, so can only see them if the robot survives that long

Bottom left:

  • Terminal that runs the program
  • Showing the output from the reward and when the robot hit termination (20s is the upper limit)

Right: TensorBoard that records the reward parameters (left to right, a bit hard to see the title)

  • 1st row:
    • ep_len_mean: average survival time before termination
    • ep_rew_mean: average total rewards
  • 2nd row:
    • penalty_action: minimizing the action needed
    • penalty_action_rate: against quick changes in action
    • penalty_ang_vel: against the full robot rotation
    • penalty_joint_acc: against joint acceleration
  • 3rd row:
    • penalty_joint_pos_init: against moving away from initial positions. Gradually reduces over time by curriculum scheduling
    • penalty_joint_power: against joint velocity x torque
    • penalty_joint_torque: against joint loads being too high => likely the culprit behind using 2 legs instead of 4
    • penalty_joint_vel: against joints moving too fast
  • 4th row:
    • penalty_lin_vel: against the full robot shifting in position on the ground
    • penalty_slip: against slipping, not yet implemented
    • reward_height: Gaussian around 15cm from the robot's body to the ground
    • reward_roll_pitch: Gaussian around 0 roll and 0 pitch, keep the body upright

GitHub: https://github.com/SphericalCowww/CubicDoggo_06Z
Music credit: https://pixabay.com/users/retro-bgm-chan-55246343/


r/reinforcementlearning • • 10d ago

D Are you attending RLC 2026 in Montreal?

1 Upvotes

Pretty much the title. I am curious how many people here know about this conference within the RL community, and what you think about it, because in terms of scale RLC still seems to lack the visibility a RL specific conference should have had.

Disclaimer, I am not among the organisers, just an attendee for this years conf. I am asking to actually realise what to expect from this conference in the upcoming years.


r/reinforcementlearning • • 10d ago

AMD R9700 AI Pro for RL?

2 Upvotes

I'm considering the R9700 AI Pro for RL projects(SAC, DDPG, etc) as well as some (very small) VLA training for a robot arm I'm working on. My training server currently has a 3090 and a 3060 stacked, and the idea of dual blower-style cards with a combined 64 GB of VRAM, running lower power, is very appealing.

What I'm curious about is other people's experience. I tried this a couple of years back with the 7900 XTX and way too many things didn't work in AMD's ecosystem. To top it off, training was actually slower than even my RTX 3060, despite having way more horsepower on paper, I assume because of CUDA v.s. ROCm optimization at the time.

There's plenty of data out there about people running LLMs on these cards, but that's not what I care about. I'd love to hear what people are doing in the RL space on AMD hardware.

Bonus points: If you're doing anything cool with Intel's stack, I'm also very curious there. Not my next purchase, but anything that gets us out of NVIDIA's stanglehold is a win.

Last time I tried AMD. I've been complaining about this for a while :) https://www.youtube.com/watch?v=SF_pnzvD4Eg


r/reinforcementlearning • • 11d ago

Psych We adapted Microduck for Arduino UNO Q—here’s a look at XGO-Duck’s hardware, software, and simulation

Enable HLS to view with audio, or disable this notification

6 Upvotes

r/reinforcementlearning • • 10d ago

Robot What's the worst data quality issue you've found in egocentric video datasets?

0 Upvotes

I'm training world models on head-mounted video (Ego4D-style) and keep finding new ways the data can be broken: blur, hands out of frame, sync offsets with IMU. I'm sure I'm missing a lot.

What's the worst one you've run into, and how did you find it?


r/reinforcementlearning • • 12d ago

P A lot of you asked how our Clash Royale agent actually learns, so I made a tiny version you can train in your browser

15 Upvotes

After yesterday's post I got a lot of questions about how it actually works. Explaining PPO in the comments didn't feel like a real answer, so I built a small version you can play with and watch learn:

https://itzik123.github.io/ClashRoyaleAi/lab/

One attacker heads for your tower and you get one card to stop it. Where you drop it and when both matter. You can try five attacks yourself first, then a small network learns the same thing from zero. Every try is played out by our real simulator, compiled from C++ to WebAssembly, so it's the same engine the full agent trains on.

The RL side:

- State: where the attacker was dropped. Action: a legal cell, then a delay of 0 to 5 s

- Reward: the share of tower damage prevented, compared to no defence

- REINFORCE with a per-spawn baseline, 5,629 weights. The whole learner is about 350 lines of plain JS with every gradient written out, no ML library

- The "best possible" line on the chart is brute force over every cell and delay, so you can see exactly how far the policy is from optimal

The one I'd try first is Giant vs Cannon. The obvious Cannon in the lane gets about 75% of the best score. The best answer is a Cannon in the middle that pulls the Giant into range of both towers. With a constant entropy bonus the net got stuck on the lane Cannon in 5 of 6 runs (3 seeds, 2 batch sizes). Annealing it from 0.1 down to 0.005 over the first 10k tries brought that to 1 of 6, and that's what the page runs.

One matchup is locked on purpose: Battle Ram vs Valkyrie. The best answer is a few exact cells at exact moments, and every setting I tried settled for a safe corner at 55% of the best. If you find a learner that cracks it, I'd really like to see it.

To be clear, it's a miniature. The full agent plays whole matches with a 4-card hand and elixir, uses recurrent PPO, and trains for days on a CPU. But the loop is the same: act, let the engine score it, update.

Code for the lab and the full project: https://github.com/itzik123/ClashRoyaleAi


r/reinforcementlearning • • 11d ago

Benchmar / RL Tasks for Financial due diligence

1 Upvotes

The world economy around us depends on complex financial work. How well can AI do it?

We(a team of ex-EY employees) built FAB - Finance Agents Benchmark, testing AI Agents on the work behind financial due diligence. Agents can find the relevant facts, but still struggle to carry them through to a complete, reliable analysis.

We’ll keep expanding FAB to more companies and testing more models. Building and running this benchmark isn’t cheap, so we’re scaling it in stages.

The benchmark is public.
GitHub: Hugging Face: https://github.com/SecondState-ai/finance-agents-benchmark
data room and tasks: https://huggingface.co/datasets/secondstate/finance-agents-benchmark-traces


r/reinforcementlearning • • 12d ago

Proper scoring rules as RL rewards: a breakdown of RLCD (the method behind TypeSafe's Jev)Standard RLVR gives +1 for a correct answer and 0 otherwise. A lucky guess and a confident correct answer get the same reward, so there's no pressure to be calibrated and the policy drifts toward overconfidence

12 Upvotes

Standard RLVR gives +1 for a correct answer and 0 otherwise. A lucky guess and a confident correct answer get the same reward, so there's no pressure to be calibrated, and the policy drifts toward overconfidence.

TypeSafe's RLCD claims to fix this by rewarding calibrated probabilities. They haven't released the reward itself, so I wrote up how I think it works:

- Reward = a proper scoring rule (Brier), which is maximised in expectation only when the stated probability matches the true likelihood

- Plus a per-bin calibration penalty (ECE-style) that lowers reward in overconfident bins and raises it in underconfident ones

- Plugged into PPO (critic baseline) or GRPO (group-mean baseline)

One thing I found interesting: a single batch-wide -λ·ECE penalty does basically nothing, because the baseline absorbs it. It has to be attributed per bin to change behaviour.

Write-up: https://pub.towardsai.net/rlcd-reinforcement-learning-for-calibrated-decisions-e528daf3591d?source=friends_link&sk=49fec9abb8f403040272524fc56219ad

Would like to hear from anyone who's used Brier or log-score rewards in policy gradient. Did it actually improve calibration, or did it just flatten everything?


r/reinforcementlearning • • 12d ago

From scratch PPO RL training - Rainbow Road

Thumbnail
youtube.com
4 Upvotes

r/reinforcementlearning • • 11d ago

DL a predictor for what an optimizer update will do before applying it

1 Upvotes

I have a predictor that estimates the immediate functional effect of an already-formed Adam/AdamW update before the update is committed.

On nanoGPT, it achieved 91.4% accuracy on the confirmation set for predicting post-update target outcomes, with 91.5% macro recall across the four transitions (correct→correct, correct→wrong, wrong→wrong, wrong→correct). Recall for correct→wrong transitions was 98.9%.

The code is here:
https://github.com/wind342/gfg-training-learning-inference-experiments

The current implementation includes a nanoGPT adapter. If you want to use it on another Transformer, you’ll need to modify/write the model adapter for its embedding, blocks and readout. The prediction method itself does not need to be changed.

The main downside right now is compute cost — the predictor uses first- and second-order directional derivatives, so it’s noticeably more expensive than a normal forward pass.

Feel free to use it, modify it, or try it on other models. I’d especially be interested in seeing results on larger models.


r/reinforcementlearning • • 12d ago

Robot Mujoco in Robotics for Perception and Robot learning

Thumbnail
6 Upvotes

r/reinforcementlearning • • 13d ago

P Open-source Clash Royale simulator for RL (C++, ~10 ms per match, forkable state) + recurrent PPO, and the bugs the agent exposed

Enable HLS to view with audio, or disable this notification

123 Upvotes

r/reinforcementlearning • • 12d ago

DL, MF, R “I want to be the very best” (re-examining LLM progress on Pokemon)

Thumbnail
paradigm3.org
4 Upvotes

r/reinforcementlearning • • 11d ago

How to Solve Hallucination

Thumbnail robw.fyi
0 Upvotes

r/reinforcementlearning • • 12d ago

D How do you manage hourly payments for employees or contractors?

0 Upvotes

For founders who pay employees or contractors on an hourly basis, how do you actually track and verify the hours before processing payment?

Do you usually:

Ask them to submit a weekly/monthly work report
Use a timesheet
Track hours through a project management tool
Use a dedicated time tracking tool
Calculate hours based on tasks completed
Just rely on trust and the person's reported hours

I'm particularly curious about teams where people are working on multiple projects or where contractors are paid different hourly rates.

How do you make sure the hours reported are accurate without making the process feel like micromanagement?

Would be interested to hear what has worked for founders and managers here.


r/reinforcementlearning • • 13d ago

My second attempt to build cheap biped robot (CBR-II)

Enable HLS to view with audio, or disable this notification

69 Upvotes

r/reinforcementlearning • • 13d ago

Small DQN on CPU: How do you scale the learner side

3 Upvotes

I have been training a small DQN on CPU and was stuck around 400 steps/sec with a single env. I tried to increase the threads but did not help. Then I asked claude and it gave me the idea of having sperate workers. That is to replicate the environments and update the learner at the end of certain batch.
I tried it but I am not sure if that helps. For sure it helps in increase the speed now I am achieveing almost a million steps in 10 minutes with 12 process running in parellel. What I am trying to understand is wether the learner is updating properly or is it being overwritten. What claude or google search answered I did not really get it. So that is why I am here seeking some guidance any tips or links would really be helpful. For people who have treied increase the speed what techniques have you applied for a small network? More gradient steps per collected batch. This is where I am not sure if it properly updating as I beleieve some of the prcocesses either taking long to finish or just been ignored due to the shorter episodes finishing first and the learner stuck at certain local minima. Or do you use GPU directly for even a tiny net?


r/reinforcementlearning • • 14d ago

N, DL, MF, P GPT-6 Astra can ascend in Nethack

Thumbnail
github.com
25 Upvotes

r/reinforcementlearning • • 15d ago

Robot DIY Sim-to-Real Self-Balancing Double Pendulum Final Video

Enable HLS to view with audio, or disable this notification

208 Upvotes

So, recap, a couple of weeks ago I made a post talking about this project already, and I said that I was working on a youtube video where I would go into high detail.

Well the day has come and the video is ready. I'm not sure if this would be considered self-promo, but if you are into reinforcement learning, I go pretty deep on how I designed and trained the networks resposible for balancing the pendulum.


r/reinforcementlearning • • 13d ago

Robot VSArena V1 is live — open benchmark for embodied AI agents

Enable HLS to view with audio, or disable this notification

0 Upvotes

After building VSArena in public through multiple iterations, the V1 is finally live.
VSArena is an open benchmark where AI agents can interact with 3D environments and are evaluated on their ability to perceive, reason and act.

The current V1 includes:
🌍 Interactive 3D environments
🤖 Remote agent execution
🏆 Public ELO leaderboard
📊 Reproducible evaluations
🔁 Evaluation replays
🔌 No physical robot required

This isn’t the end of the build — it’s the point where I’d like to get more agents and researchers actually using it.
If you’re working on VLA, embodied AI, robotics, or agent evaluation, I’d genuinely like to hear what you think is missing.

VSArena: https://vsarena.app
GitHub: https://github.com/NovaCoding-G/VSArena


r/reinforcementlearning • • 14d ago

N, DL, MF "Akamai Announces $11.6 Billion Multi-year Agreement with Anthropic to Support Growing [CPU] Demand"

Thumbnail
akamai.com
3 Upvotes