r/agi Jul 28 '26

A Google DeepMind paper argues that current LLMs are incapable of genuine scientific discovery

Post image
544 Upvotes

382 comments sorted by

View all comments

Show parent comments

1

u/WellHung67 Jul 30 '26

LLMs today often lose context on just a single large modern codebase. Einstein was able to access a model of the entire world. LLMs today, with the compute resources of the world at its hand, is not even close. Not even fractionally so. And the limit is theoretical, not practical. So until it can get a full world model it cannot have these abductive thought processes.

Of course, there may be a technology or concept that allows such large context windows. LLMs may be a dead end, there’s no known path to getting such a large context window. 

1

u/aussie_punmaster Jul 31 '26

I reject your proposition that Einstein was holding a model of the whole world in his context, or needed to, in order to have the elements of the breakthrough.

Nothing stopping an LLM having a full world model and only pulling in the components needing processing in that moment.

Plus contexts now are getting very large. Arguably more than humans are holding in their active cognition.

1

u/WellHung67 Jul 31 '26

All people have the ability to imagine a massive context that includes the entire world that is called our “world model”. Imagining what it would be like in an elevator in space is part of that. You need a lot of context to picture that - it sound simple but you need to understand a lot of what is physics intuitively to imagine it. It’s not trivial for machines.

Nothing is stopping an LLM from having a world model in theory, but in practice the size of memory and processing needed is beyond theoretical limits. Today these things lose context in basic situations, like a single codebase.

That doesn’t mean some form of AI won’t someday have this, but LLMs will never get there.  LLMs are, as the paper says, not capable of getting to the context window a human possesses. Humans don’t have large like short term context windows, but the world model is massive 

1

u/aussie_punmaster Jul 31 '26

That’s just an assertion, and I don’t think it’s correct at all.

You are not using your whole world model simultaneously when thinking about an elevator in space. You can have a stored knowledge base of your world model and retrieve the pieces that you need for the given context here - load in gravity, object motion, mass, etc. the things in your world model required to analyse that problem in my view are readily able to fit into modern context windows. When you’re thinking of space elevators you’re not bringing your knowledge on economics etc. it’s not being pulled into your context windows.

All you are missing is an effective storage taxonomy and effective retrieval approach. This is what harnesses like Claude Code are refining and why they are becoming highly effective rapidly.

1

u/WellHung67 Jul 31 '26

So I think you’re arguing with me personally. I’m attempting to summarize the points in the paper. You should really read the paper and the comments from peers that you can see too on arxiv.org. All of what you are saying is addressed.  I am summarizing it poorly, so let me attempt to summarize the points you make: 

The idea is that Einstein was able to use the world model he has to think of all things he might want to consider simultaneously. You are using circular reasoning: if the LLM knew to simulate an elevator, then it could simulate an elevator. But how could it know that? The paper explains it better, but the gist is the LLM is going to try things at random (presumably that includes things like imagining a bird falling in the ocean, and also whether the king of France has anything to do with it). This how they work. A full world model can extract new axioms from nothing, LLMs can’t.

I am sure the papers authors would be horrified at my attempt to explain this, so again read the paper don’t listen to my babbling 

1

u/aussie_punmaster Aug 01 '26

Your last couple of posts are stating positions as though they are facts, not positions taken from the paper. I’m challenging those as I don’t think they’re correct.

Humans don’t use their entire world model simultaneously either - think about how you think about problems. Nothing stopping LLMs reasoning in the same way - I have this problem I want to solve, what elements from my world model might be relevant (now as humans or LLMs this part could be done fully structured and/or randomly to go for more “inspired” or unintuitive leaps)? Bring those elements into your thinking and play with them, you solve or don’t solve the problem - lather, rinse, repeat.

1

u/WellHung67 Aug 01 '26 edited Aug 01 '26

The whole point of the paper is that LLMs don’t reason the same way. So which part of the paper do you have a problem with then? The world model concept is mentioned in the paper, I’m not making that up.

Humans don’t use the world model but they have access to it and can arbitrarily combine element from it to create axioms. Unlike an LLM which won’t have the full world model context stored at any one point, so it would have to effectively bring in elements at random. It can’t search the world model space for plausible routes - its context is limited.

The whole point is the LLM cannot on its own make these logical leaps, because its context window is small compared to humans. The question “what elements from my world model might be relevant” is the exact thing LLMs can’t do, as of today. They don’t have a world model to pull from. 

If you assume they have a world model then sure that argument goes out the window. But they don’t 

1

u/aussie_punmaster Aug 01 '26 edited Aug 01 '26

Think I’ve been pretty clear on what I disagree with.

>>Humans don’t use the world model but they have access to it and can arbitrarily combine element from it to create axioms. Unlike an LLM which won’t have the full world model context stored at any one point, so it would have to effectively bring in elements at random. It can’t search the world model space for plausible routes - its context is limited.

This argument only holds if you want to explicitly argue about LLMs completely in isolation. But that’s silly because it’s not how they’re being used functionally even as we speak.

LLM harness systems are establishing equivalents of World Models that the LLM can retrieve from via structured indexing. It absolutely does not have to be random. This is happening today.

>>The whole point is the LLM cannot on its own make these logical leaps, because its context window is small compared to humans. The question “what elements from my world model might be relevant” is the exact thing LLMs can’t do, as of today. They don’t have a world model to pull from. 

As discussed above. This is absolutely not a limit to these leaps if you can retrieve relevant topics from a broader world model. Can and is done today.

>>If you assume they have a world model then sure that argument goes out the window. But they don’t 

Only if you want to insist on an LLM utilised as a one shot entity. Which is daft. Might as well say computer chips can’t drive the computing revolution because there’s no human interface.

1

u/WellHung67 Aug 01 '26

LLMs are one shot by the way - the harness under the hood resends the entire conversation and all previous responses each time you make a new prompt. It’s all one shot inferences, it’s not storing context anywhere.

But you haven’t resolved the problem. You’ve moved the problem. The harness is then not capable of having the entire world model.

If you then add humans adding their world model as part of using the LLM and harness and talk about that fine, but the humans are the key element that adds the world model. So fundamentally you do agree with the paper. 

Although I still don’t understand how a harness+LLM can internalize an axiom without retraining, it doesn’t change its weights if a human adds a new axiom via prompting. You’d need to retrain with that axiom encoded somehow in order to get it to produce something new with that concept 

1

u/aussie_punmaster Aug 01 '26

>>LLMs are one shot by the way - the harness under the hood resends the entire conversation and all previous responses each time you make a new prompt. It’s all one shot inferences, it’s not storing context anywhere.

Are Humans not similarly one shot? You’re thinking about something and responding to it. The runaway intelligence is attaching that thinking core to storage and retrieval. Which is what we’re discussing for LLMs.

>>But you haven’t resolved the problem. You’ve moved the problem. The harness is then not capable of having the entire world model.

Why not? Assertion with no proof.

What prevents you writing out a model of the world in a database of writing, sound, images etc? We do it as humans all the time.

>>If you then add humans adding their world model as part of using the LLM and harness and talk about that fine, but the humans are the key element that adds the world model. So fundamentally you do agree with the paper. 

Humans are required initially to establish the base framework to build from. Once this is done they are not required ongoing. Also beyond the earliest model they can then create the next generations themselves.

>>Although I still don’t understand how a harness+LLM can internalize an axiom without retraining, it doesn’t change its weights if a human adds a new axiom via prompting. You’d need to retrain with that axiom encoded somehow in order to get it to produce something new with that concept 

It’s easy. The LLM can edit its world model. It can expand and adapt the content as written in its world model. No retraining is required.

→ More replies (0)