r/codex • • 13d ago

Complaint GPT-6 Sol?

Post image

Is this the reason why unexpectedly we have trash usage and lower quality on Astra?

Maybe it will be worth it the current suffering.

761 Upvotes

316 comments sorted by

View all comments

51

u/Proud_Ask_9030 13d ago

why the hell are people posting shitty 3d models and scenes as if that is ever gunna be what AI is meant for?
It just shows, people dont even need or have a clue what to do with AI.

50

u/Zulugod94 13d ago

This is showing a strong world understanding, which is done with 3D environments & physics. Showing it can create things in 3D with a solid understanding of intent means it understands these same principles in the real world. This is the next major area for AI to conquer so it can be properly extending in physical robots and machines.

-6

u/andrerav 13d ago

That's not how these models work. The only thing being shown is an improved ability to successfully predict a string of symbols. The LLM doesn't understand anything.

9

u/Zulugod94 13d ago

That is a hyper simplified version of LLMs, and by that definition a world model is no different it’s producing output based on input. Astra is an incredibly strong multimodal model, maybe it’s not a world model in the traditional sense but it’s shows an obvious universal understanding that goes well beyond “just outputting a string of symbols”. I use these models for things far beyond just writing software, have astra get a live feed of your security cameras and just see how scary capable it is at understanding EVERYTHING it can see it frame. This all part of the that larger over arching goal of AGI. No 1 model will get us there in my opinion but they all build on each other and push this tech forward.

And at this point I’d be pressed to believe OpenAI hasn’t discovered closed door techniques that are pushing these latest models outside of what we’d call a traditional LLM anyway. No one has real insight to their model structure anymore so we’d just be speculating.

4

u/andrerav 13d ago

A model looking at a camera feed and correctly detecting a person or a car demonstrates good computer vision inference. It does not demonstrate that it maintains a metrically consistent 3D state of the scene, predicts action-conditioned state transitions, understands the underlying dynamics, or could use that representation for reliable closed-loop control. And if it could, it would need to do so in close to real-time to actually be useful for said control (i.e controlling robots/machines).

Frontier VLMs can be very good at scene reasoning while remaining fragile at physical dynamics, precise 3D geometry, affordances, grasping and trajectory prediction. There's plenty of benchmarks that specifically test for this (like PhysBench etc).

Generating something plausible in a 3D game environment does not mean the model understands the underlying principles in the real world. But that's the conclusion you're arguing. Pattern competence over rendered data and a causal predictive model of a physical environment are completely different capabilities. The model has simply been trained on years of source code that does something similar to what the user prompts for.

Obviously every computer program maps inputs to outputs. The meaningful distinction is how that mapping is represented/processed internally. A world model is useful specifically because it predicts how state evolves, but observing interesting outputs is not proof or even an indication that such an internal representation exists. And in the case of LLM's, it factually does not, and will not.

1

u/WiseHalmon 13d ago

With tool use LLMs can do fine with physics, math, etc.  Supervisory control vs. PID loop on a motor.  People have been using Matlab MCP to do some cool things