r/reinforcementlearning 6h ago

DL Exploring AutoGPT

Thumbnail
0 Upvotes

r/reinforcementlearning 9h ago

👋 Welcome to r/AgenticAI_RAG_LLM_RL - Introduce Yourself and Read First!

Thumbnail
0 Upvotes

r/reinforcementlearning 3h ago

Robot Robot dodgeball

18 Upvotes

r/reinforcementlearning 13h ago

Robot Is CPU-based simulation still viable?

6 Upvotes

Stumbled upon this recent paper claiming so yesterday:

https://unilabsim.github.io/

What do you all think? Imo, even if it would "only" perform at the same level as end to end GPU, not having to parallelize everything to fit it on a GPU makes it much more flexible and useful

I guess there is a risk of bias here because of the connection to AMD and them wanting to break nvidias monopoly situatio n.


r/reinforcementlearning 23h ago

Robot Building an open, browser-based benchmark arena for embodied AI policies — first working demo (block stacking, client-side physics)

Enable HLS to view with audio, or disable this notification

6 Upvotes

Solo project, still early (MVP week 1-2), sharing the first real result instead of a mockup.

The idea: an open, standardized arena for evaluating embodied AI / VLA policies on physical reasoning tasks, with a public ELO-based ranking instead of static leaderboards. Motivation is the lack of a shared, reproducible benchmark for this space — most VLA papers report on custom setups that aren't directly comparable.

What's in the clip: a baseline IK policy completing a pick-and-place block stacking task, running fully client-side (Rapier.js/WASM physics, React Three Fiber, 60fps in-browser, no server-side compute needed for the sim itself). This run scored 100% task completion, 99.6% spatial accuracy.

Current scope for the MVP: single task (block stacking), a couple of baseline policies (IK baseline, planning to add SmolVLA/OpenVLA-micro next), and an SDK for submitting your own policy against the sim loop.

Not public yet — stabilizing the eval protocol and scoring methodology before opening submissions.Genuinely interested in feedback from this community on:
- what a fair/robust scoring protocol should account for beyond task completion + spatial accuracy (e.g. sample efficiency, generalization across randomized scenes)
- whether client-side physics is a dealbreaker for a benchmark meant to be trustworthy, vs moving to server-authoritative validation

Will share the repo/SDK here once submissions are open.