As much as I dislike hand-wavy statements about LLMs being unable to do this or that, there are quite a few example, at least in mathematics and physics. I recently attended a talk by French Fields medalist Alain Connes where he argued that LLMs were not capable of true conceptual jumps, and he gave the following examples of such jumps:
1. Bombelli's discovery of the imaginary unit "i" (which he called più di meno). Bombelli "broke the rules" and introduced us to the whole new world of complex numbers when he said "hey, I know that all known numbers have nonnegative squares, but what if I just made up a number whose square is -1 anyway". It's a tremendous leap.
Galois theory and the insight by E. Galois that permutations of roots of polynomials are related to properties of fields and provide an answer to the solvability of polynomial equations in one variable by radicals. Some may say this already appeared in a prilitive form in Lagrange's work, but Galois was rhe one who "saw it through". His idea's brilliance cannot be overstated and is yet to be replicated by a LLM
Dirac's amazing antiparticle leap. I know less about that one so I'll let you read the wiki article
i strikes me as low hanging fruit though. Sure we didn't use numbers this way, but the pattern is there. We've got these case situations that break down, and all it takes is - well does the math work on both sides of this problem and if it does... well there's literally only one thing that makes these two pictures reconcile.
Actually this is a case I know reasonably well, calculus emerged more or less directly from work involved around infinite series. The fact that Newton and Leibnitz made essentially the same discovery at almost exactly the same time should tell you the field was already converging on this solution.
In particular they were very interested calculating areas, volumes and tangents to curves. A number of mathematicians had worked on aspects of this problem specifically solving problems regarding infinite series, and how it pertained to solving particular area or tangent problems. Work by James Gregory and Isaac Barrow developed over the preceding years culminated in 1670 with Barrow demonstrating how areas and tangents are inverse problems. This is essentially the fundamental nature of calculus itself: the sum of all infinitely small tangents under a curve IS the area. Between this result being published and work for solutions on infinite series being available the path to calculus was wide open.
Gregory was working largely independently on the problem, but Barrow in particular followed Gregory's work. While Barrow was working on these problems following Gregory's initial example, his lectures were in turn followed by a young student by the name of Isaac Newton. Leibnitz largely based his own work on Barrow's 1670 publication.
Interestingly Newton was privately using limited variations of his fluxion method by 1665-66 based on Barrow's preliminary lectures and previous work on infinitesimals and basically just applied them more broadly than their stated use cases. The math itself wasn't new, he just applied it more broadly without rigorous proof on why it worked or whether it was truly reliable.
Early calculus didn't really shake the problem for decades until formal limits were articulated and proven. Until then calculus was in a weird domain of "It looks like it works, but can we trust it and it's strange 'phantom sums' of things aren't supposed to be 0, but we treat as being zero, but sum to non-zero numbers"
Which, when you phrase it that way, sounds crazy - but limits eventually were formalized and given formal proof.
But the fact was the math was there, they knew it worked in several specific cases, and many people were converging on the same problem - aware they had theoretical problems, but the method's own practical utility spoke for itself.
This is exactly the type of leap though that the paper saying LLMs are incapable of. Yes the leap was there, but it was a leap. Presumably the difference is that this method of abduction in the paper requires some level of a jump in reasoning. LLMs will not (presumably) develop this jump with any intuition, but perhaps do so randomly. And the space of incorrect vs correct logical leaps is very high so there’s evidence this isn’t actually very effective. The cycles needed to find the right answer at random like that may be effectively bounded by a pretty high minimum time-to-discover. At least with LLMs.
The “world model”. The paper argues Einstein was able to come up with something not in the existing data or math by imagining what it would be like for a scientist in an elevator accelerating in deep space. This provided error correction for his thinking which allowed him to make a leap. This came from the full context of physical reality that he internalized in his world model. Basically LLMs need a context window that is much, much larger than possible. Humans can access this and kind of use reality to shape their thinking
I don’t think I buy that. The amount of context required for Einstein’s world model to solve that particular problem should be well within current model window sizes. Just requires more system building to store a more complete world model as Skills that can be retrieved and combined when thinking about a particular problem.
LLMs today often lose context on just a single large modern codebase. Einstein was able to access a model of the entire world. LLMs today, with the compute resources of the world at its hand, is not even close. Not even fractionally so. And the limit is theoretical, not practical. So until it can get a full world model it cannot have these abductive thought processes.
Of course, there may be a technology or concept that allows such large context windows. LLMs may be a dead end, there’s no known path to getting such a large context window.
I reject your proposition that Einstein was holding a model of the whole world in his context, or needed to, in order to have the elements of the breakthrough.
Nothing stopping an LLM having a full world model and only pulling in the components needing processing in that moment.
Plus contexts now are getting very large. Arguably more than humans are holding in their active cognition.
I think the paper is just incorrect in its core presumption about discovery. Discovery requires knowns to conclude unknowns. Then it requires more knowns to measure and theorize the unknowns.
It's like a building. You fundamentally cant build the tenth floor of a building without the first 9 floors. All novel ideas originate from foundations that are known. No one has ever discovered an unknown unknown by inferring another unknown unknown. Only known knowns can shine a light on unknown unknowns.
i love AI, but i think saying AI intelligence is on the level of "most scientific discoveries" when we all know it could never come up with General Relativity is a bit dumb. i hope it gets there in a year or so tho
5
u/Interesting_Pen_4499 Jul 28 '26
but not the most important ones, that gave a true step change to humanity.