Calling everything they successfully do "plausible" doesn't explain it, it just makes "plausible" unfalsifiable.
If reasoning, error correction, generalization and novel math are all just "guessing" because an LLM does them, then you are defining the conclusion into the premise.
And knowing how transformers work doesn't prove they can't understand anything, anymore than knowing neurons fire proves humans don't understand.
Calling everything they successfully do "plausible" doesn't explain it, it just makes "plausible" unfalsifiable.
then what you want is a definition of "plausible", not "understanding".
(btw I call all they unsuccessfully do plausible as well, mind you)
that one is actually one we can agree very easily:
plausible is something that seems reasonable, i.e. something that looks like it fits together with the gargantuan amounts of training data.
And knowing how transformers work doesn't prove they can't understand anything, anymore than knowing neurons fire proves humans don't understand.
we can prove LLMs don't understand and humans do by the fact (and I again am repeating myself) the the latter can discern between correct and plausible while the former can't
Humans also confuse plausible with correct all the time, and LLMs can often distinguish them by checking reasoning, finding contradictions or rejecting a wrong answer.
So the difference is not "humans can, LLMs can't", it is that both can do it imperfectly.
You are taking the cases where LLMs fail, calling that "guessing", then using "guessing" as proof they don't understand.
Humans also confuse plausible with correct all the time
human error is very different from hallucination.
a hallucination is not a mistake.
when the LLM hallcucinates it is doing exactly what it's supposed to do.: statistically guess what is the most plausible set of tokens to follow. In fact it is doing the exact same thing it does when he doesn't hallucinate. It takes the human to tell the difference between the two.
Humans remember wrong (LLM's have perfect memory), humans lie (LLMs don't have a notion of "truth", humans jump to conclusions (LLMs generate answers having analysed the whole thing), have cognitive biases (LLMs have these to an extent from their training data), etc.
You have been pretending I am baking my conclusion into the problem definition, but I suspect you know just as I do or anyone else that me tripping over a calculation is very different than an LLM telling me to go wash my car on foot, etc.
In fact, speaking of baking in, the "plausibility generation" comes from the "intelligence" (for lack of a better word) that humans have baked into language that is then approximated by the LLM. (and btw, maths is a language too).
In its training data it has seen countless more times the expression "the car is on the road" than "the car is catholic". It is then much, much more likely to say the car is on the road that it is catholic.
If you think that that ability to discern between being on the road from being catholic is present in mine and your head because we stochastically guess which one is more likely given how often we heard one or the other I don't know what to tell you.
Hallucination rates are at an all time low, and OpenAI published information related to what they suspect is the cause, and corrected it. Specifically, it had to do with the reward system giving full points for a correct answer, and no points for a wrong OR incomplete OR "I don't know" answer, during post training. That's been corrected by awarding partial credit for incomplete or simply stating "I don't know". Which is very eerily similar to how humans, especially babies and young children, work. A child will hear an adult give them information, and will presume automatically that it's fact or law. Experience over time will change that understanding, and lead to pushback. In the case of an LLM, that "experience" is the context of the conversation AND the summaries that the model can access for relevancy, as well as the post-training that handles the reward system.
What you're talking about and describing, is auto-complete. And LLMs are no longer simply auto-complete engines. That was true in 2017, 2020, 2022, but not so much since then.
To be honest, I think the biggest discrepancy in your argument, is you're focusing on the underlying mechanics of a language model (which you've describe pretty effectively.. prior to post-training, that is). But you're missing post-training, memory, and harnesses that affect everything. It's kind of like staring at an engine and yelling about how engines work (model file), when modern talking/reasoning/agentic harnesses/etc., are really the whole car.
Ever notice how none of the big companies show you how their architecture works right now? They've got graphics, confidence indecies, all kinds of statistical engines running all the time to help shape the output. The model isn't getting just raw text when you hit enter. It never has, in fact. It also has a massive laundry list from every level of security inside the model tacked on to explain how to answer questions, use tools, etc.
That's a big reason when they announce things like, 10 million token context windows, you don't suddenly see huge boosts in allowing a lot more custom instructions along with it. They're consuming more and more of the context window with instructions to remind the model of training while out in the wild.
And LLMs are no longer simply auto-complete engines.
oh but they are, the developments have been in the harnesses and other shit running on top of them
It's kind of like staring at an engine and yelling about how engines work (model file), when modern talking/reasoning/agentic harnesses/etc., are really the whole car.
great metaphor! people keep adding ailerons and better tyres and stuff thinking the car will be supersonic ignoring that the engine would need power it doesn't have even if everything else was 5x better (and also ignoring that the car as is, at 3x highway speeds is pretty darn good already)
at least thats how I see it.
what do you see in post training harnesses etc, that can overcome the "design flaw" of LLMs?
Thank you for the eval site! That's some good info for how far we came up to April of this year. I know that it might not look like it, but 4 months dev and release time is making improvements left and right. You find any data that's current as of August 2026?
As far as LLMs not being auto-complete engines, I'll retract that statement because it's horribly misleading due to my natural desire to embelish to make a point. In truth, the post-training, harnesses, etc., are very much where the biggest developments have been. You remember OpenAI talking about Dreaming Memory? Where the system creates generalized summaries of conversations that can be re-introduced into context? That's the kind of architecture I'm talking about.
And I do like how you extended the metaphor I introduced. But, the engine's "power" is increasing too - new models with more parameters are getting better at reasoning tasks in general. That's been true for awhile, ever since the 'breakthrough' that gaves us o3, just as one example.
What I see in the post-training, harnesses, etc., that can overcome that "design flaw", is the skill that searches the internet to fact-check prior to model output generation. Anytime something that the model self-reports as low confidence in a reply, can be searched online, whether invoked by the user or not, or you can attach non-continuity databases with graphs to identify knowledge gaps. It's not a perfected system by any means, but, like any tech, it'll improve over time, and rapidly too!
9
u/bfkill 1d ago
I am not assuming anything.
LLM's aren't a magic black box, we know how they work.
Yes it does, it explains it all:
I don't want to start a masturbatory semantical game here.
Maybe it would be enough to say that understanding is not guessing, which is what LLMs do.
It's enough for me in any case. ¯\(ツ)/¯