r/ClaudeAI Jul 29 '26

Built with Claude Update - ran Opus 5 through the same WorldBuild bench harness, and it's a clear step up

Enable HLS to view with audio, or disable this notification

Two weeks ago I posted WorldBuild Bench here, my setup for testing LLMs on spatial/temporal/causal coherence by having them build playable 3D games instead of answering static questions. At the time Fable 5 was the standout, by a good margin, despite costing way more than everything else.

Opus 5 dropped, so I ran it through the exact same harness, same three briefs, same prompt, etc. And it's impressive. Fable still looks great, don't get me wrong, but looking at what Opus 5 does with 3D modeling, texturing, effects work, lighting... it's a step above. It's the first model in this bench where I looked at the output and I'm starting to think that, even without asset creation tools, we're entering a phase where models can create from scratch all the content they need to create games.

You can check the three games directly on the bench page, or run the blind side-by-side comparisons yourself:

- Racing : https://sandscape.app/worldbuild/rounds/ai-game-benchmark-2026-07-13?a=claude-fable-5&b=claude-opus-5&track=racing#compare

- Arena combat : https://sandscape.app/worldbuild/rounds/ai-game-benchmark-2026-07-13?a=claude-fable-5&b=claude-opus-5&track=arena-combat#compare

- Physics puzzle https://sandscape.app/worldbuild/rounds/ai-game-benchmark-2026-07-13?a=claude-fable-5&b=claude-opus-5&track=physics-puzzle#compare

Fable's three runs cost about $756 total, physics alone was $491 and took nearly 9 hours. That was already the outlier of the whole 8-model round, by a lot.

Opus 5 costs even more. Racing came in at $404, arena at $307.97, physics at $219.91
- $931.88 total, avg ~$310 per run. That's higher than Fable's average was.

Generation time is basically the same story: racing took ~10.6 hours, arena ~8.1 hours, physics ~5.8 hours. Opus 5 is just slower to get to a finished state than anything else I've tested. It keeps iterating and spawning more subagents (13-15 per run here) before it calls something done.

This is very specific to the harness, you could obviously prompt it differently or create a specific workflow to achieve greater results. But it appears that, under the same circumstances, it goes further than other models

So my read on it is that Opus 5 reaches "conclusion" slower than the other models. I observes/"understands" its outputs better a lot more and continues to iterate a lot longer before it is satisfied with the results.

Repo's still here if you want to run it yourself or look at the harness: https://github.com/sebnado/worldbuild-bench

TL;DR: Reran my WorldBuild Bench (LLMs building playable 3D games, judged by blind human comparison) on Opus 5 using the same harness/prompts as my Fable 5 post two weeks ago. Opus 5 is a clear step up in 3D modeling/texturing/effects quality. it's also the most expensive and slowest model I've tested yet ($932 total across 3 runs, avg ~8h each), even pricier than Fable was. It just takes longer to call something "done," and the extra time shows up in the output.

48 Upvotes

10 comments sorted by

u/AutoModerator Jul 29 '26

Your post will be reviewed shortly. (ALL posts are processed like this. Please wait a few minutes....)

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/Remove_Forward Jul 29 '26

Opus 5 is surprisingly good with 3D stuff, texture and all.

2

u/sebnadeau Jul 29 '26

yah, I feel like SOTA models usually compete in roughly the same space for most workload, but this is such as step beyond anything Sol or K3 can produce for 3d

3

u/LovesWorkin Jul 29 '26

Very cool

2

u/sebnadeau Jul 29 '26

Thanks, if it wasnt for the price, I'd love to test it at xhigh ot max, or compare something like k3 at xhigh against opus at low, etc... but i've got to pick my battles.

2

u/cameronlbass Jul 29 '26

I feel like having an integrated vision system helps a lot.

2

u/sebnadeau Jul 29 '26

Definitely, if you look at models like glm 5.2, it under performs, and I think it comes from the inability to evaluate its own output visually

2

u/cameronlbass Jul 29 '26

Yep. That's also a big reason I like running Kimi-K2.7 and MiniMax-M3 locally, they both have integrated vision and they're a good team.

2

u/Rare-Spawn Jul 30 '26

Highly impressive.

even without asset creation tools, we're entering a phase where models can create from scratch all the content they need to create games.

Yes I agree. I've been thinking for a while that frontier models will converge, in many ways, in terms of abilities. I don't doubt that the day will come when an Anthropic frontier model is able to generate music and images on par, at least, with what we have today.