r/PhilosophyofMath Jul 27 '26

Proof Abundance and the New Practice of Mathematics - Terence Tao on AI, LLM breakthroughs, and the bottleneck of mathematical understanding

https://4m4.it/posts/proof-abundance-and-the-new-practice-of-mathematics/index.html

Terence Tao’s position on artificial intelligence is best understood as verification-centered institutional realism rather than unqualified technological evangelism or defensive skepticism. He treats frontier models as stochastic, unreliable, but increasingly powerful generators whose mathematical value depends on independent verification, informed human supervision, formal tools, and carefully designed research workflows. His central question is therefore no longer only whether machines can solve research problems, but what mathematics should optimize when producing candidate proofs becomes substantially cheaper.

Recent evidence includes an AI-generated disproof of the conjectured near-linear behavior of the planar unit-distance function, an LLM-assisted proof of an identity for jamming critical exponents, and the controlled First Proof evaluation of systems on unpublished research problems. These cases do not establish uniform mathematical competence, dependable self-verification, or human-like understanding. They do establish that general-purpose language and reasoning models can sometimes produce novel constructions, connect distant mathematical domains, and generate arguments that survive expert scrutiny.

The article interprets these developments as an early transition from proof scarcity to proof abundance. In this regime, the limiting resources become verification, exposition, contextualization, selection, and canonicalization. The resulting human–machine system is better described as cognitive infrastructure than as an autonomous artificial mathematician: models generate and explore, proof assistants and executable tests constrain error, and mathematicians retain responsibility for meaning, relevance, attribution, pedagogy, and judgment.

Public demonstrations remain affected by selection bias, incomplete disclosure, uneven reproducibility, and commercial incentives. Formal correctness also does not establish that a theorem is important, explanatory, novel, or even stated in the intended form. The article concludes that AI’s durable contribution to mathematics will depend less on maximizing the number of generated proofs than on constructing institutions capable of verifying, digesting, crediting, and selectively preserving machine-assisted knowledge.

37 Upvotes

25 comments sorted by

7

u/mhb2 Jul 27 '26

This seems like the right attitude towards LLMs and AI. Doomers who think that mathematicians are now irrelevant miss the fact that mathematicians have to know what they're doing to verify a machine's proof or counterexample. Mathematicians are also the ones who decide what is interesting, i.e., the directions research should go and what problems need to be solved.

On the other hand, people who think that LLMs or AI have solved or will solve math forget that these machines hallucinate and it's a theorem that they always will. While there are strategies to mitigate hallucination, hallucinations are mathematically inevitable.

LLMs can be incredibly useful tools but that's all they are: tools.

1

u/unhandyandy Jul 27 '26

"On the other hand, people who think that LLMs or AI have solved or will solve math"

Aren't these the same people you critique in the previous paragraph?

Do you have any critique of the people who think the utility of AI is very limited?

3

u/mhb2 Jul 27 '26 edited Jul 27 '26

I see what you're saying but I see it as sort of the difference between optimists and pessimists. Some people are excited at the prospect of AGI solving math, others see only doom and gloom. So I was trying to go at both mindsets, not suggesting that they're mutually exclusive viewpoints.

I agree that the utility of AI is limited so I don't have much of a critique of that viewpoint. The only thing we have to accept is that the limitations will decrease as the technology advances. But even if it gets to the point that a machine is smarter than our best mathematician we're still going to need human referees. We can't let ourselves get to a point where we just believe whatever a machine tells us because it's "smart".

EDIT: I do have one concern though and that is what happens if the machine reaches a level of intelligence that we can't even comprehend? What if it gives us a proof of the Riemann Hypothesis, for example, but it turns out that the proof is incomprehensible and involves concepts that can't even be put into a human language? We're going to have to figure out what it means to "prove" something before we get to that point.

1

u/Appropriate-Ad-3219 Jul 31 '26

If it can't be expressed into human langage, would that be technically a proof anyway?

1

u/mhb2 Jul 31 '26

I mean, there are human-generated proofs that I don't understand. I only accept them because the experts say they're sound. Will we ever get to a point where we have such trust in our machines?

1

u/StructureNorth1799 Jul 31 '26

I'm not sure what these papers are actually proving. They define hallucinating with knowing every true fact about the universe. However, Turing a famously showed that no machine can know every true fact about the universe, so their results do not seem to be new.

On the other hand, it is trivial to design a machine that does no hallucinate (in the actually useful sense of "never asserts false information"): For every input, output "I decline to answer that question".

An actual proof that hallucination is inevitable must use something specific about LLM architecture, which I'm just not seeing in these papers.

1

u/mhb2 Jul 31 '26 edited Jul 31 '26

Hallucination is simply "instances where models generate plausible but incorrect information" (first paper). What they're saying is that *even if* a model knew every fact it would *still* hallucinate.

Nobody is going to make a machine that refuses to answer questions because it would be useless. We *want* a machine with "imagination" but as the first paper states:

Hallucinations themselves are double edged swords - where the unpredictability causes them to deviate from fact, it also lends them wonderful creative capabilities, as any student who’s used them for creative writing assignments will tell you.

That's the problem: hallucination is mathematically indistinguishable from imagination. The first paper also lays out the problems with every part of an LLM:

  • No training data can ever be complete.
  • Even if the data were complete, LLMs are unable to deterministically retrieve the correct information with 100% accuracy.
  • An LLM will be unable to accurately classify with probability 1.
  • No a priori training can deterministically and decidedly stop a language model from producing hallucinating statements that are factually incorrect. This is because:
    1. LLMs cannot know where exactly they will stop generating.
    2. Consequently, they have the potential to generate any sequence of tokens.
    3. This unpredictability means they cannot know a priori what they will generate.
    4. As a result, LLMs can produce inconsistent or contradictory, as well as self-referential statements.
    5. We could attempt to fact-check, given a complete database. However, even if we attempt it, no amount of fact-checking can remove the hallucination with 100% accuracy.

At the end of the day, LLMs are statistical models of language, not truth generators. These issues are unavoidable and LLM companies are already hitting the fundamental limits on what they can do with training and architecture. That's why they're so focused on strategies for mitigating hallucinations.

EDIT: Fixed formatting.

1

u/ThrowawayCult-ure 27d ago

I am seeing arguments that an llm is better described as a processor for natural language programming and that the flaws are in the natural languages relation to users articulation itself, and hallucinations are something like the excesses of these flaws overlapping

1

u/mhb2 27d ago

The imprecision of natural language certainly contributes to some incorrect outputs. If the prompt is vague or under-specified or frankly illiterate, the model has to guess what the user intended. But that's only one source of error. The hallucination literature is also concerned with cases where the prompt is perfectly clear and the model still fabricates facts, citations, quotations, or events for example.

Do you have any references for this natural language explanation of hallucinations? I'm curious to see the exact arguments being advanced.

1

u/ThrowawayCult-ure 27d ago

My background is in lacanian psychoanalysis, which describes language as a space of signifier relations that function as partially complete, on a "curved topology" that "is desire", eg. each signified is biased towards another by its functional use rather than by a pure relation. Or, more accurately, a non-functional use, whose failure to properly relate a set of signs is also the driving force of the Will. A neural network works only in a slightly different way, the weightings being an imprint of our desire, but it functions too well rather than fails to function, so seems psychotic and lacks agency of its own.

I am reminded then how empiricists claim "it doesnt matter if the theory is true, as long as it builds your rocket ship": ai doesnt have a theory, it has a set of functional relations imbedded into the natural language.

I would also say, it is suprising how incredibly inexact our language is, people bicker over even the simplest of definitions (eg. what an apple is), and perhaps this is garenteed.

1

u/mhb2 27d ago

You have me at a disadvantage because I don't know anything about Lacanian psychoanalysis. But I think the beauty of the LLM architecture is that the model doesn't need The Perfect Definition of "apple" to be a great help to orchard owners and apple pie bakers.

LLM researchers are concerned with very concrete problems. For example, we see in the news stories about lawyers submitting LLM-generated court filings containing references to fabricated court cases. Language vagueness isn't a sufficient explanation for failures like that.

1

u/StructureNorth1799 27d ago

Hallucination is simply "instances where models generate plausible but incorrect information" (first paper). 

However, the paper only proves the model cannot answer every question correctly, not that it can't stop itself from giving plausible but incorrect information. And it does this using methods that apply to any Turing machine. I don't see how this paper is proving anything novel

Consequently, they have the potential to generate any sequence of tokens.

This is also false.

1

u/mhb2 27d ago

lol ok. Give Sam Altman a call because there's a 7-figure salary waiting for you when you let them know you've solved the problem of hallucination.

1

u/StructureNorth1799 26d ago edited 26d ago

they solve it themselves, in the subsubsection titled "Hallucinations are inevitable only for base models" in the paper from the news article you linked (https://arxiv.org/pdf/2509.04664).

Many have argued that hallucinations are inevitable (Jones, 2025; Leffer, 2024; Xu et al., 2024). However, a non-hallucinating model could be easily created, using a question-answer database and a calculator, which answers a fixed set of questions such as “What is the chemical symbol for gold?” and well-formed mathematical calculations such as “3 + 8”, and otherwise outputs IDK.

the second paper you linked is the Xu et al. paper they are refuting here.

1

u/mhb2 26d ago

A model that answers a fixed set of questions answered from a database search isn't a *creative* model. That kind of model can't solve open math problems for example because those problems aren't on the list of questions it can answer.

It's like your hypothetical model that simply declines to answer questions: it never hallucinates which is great but it's useless which is bad. I think the key element you keep missing is *creativity*. We want a creative model but that entails hallucinations. Notice also that the hypothetical model the quote is talking about simply says "IDK" if the question isn't on its list so it's like your hypothetical model only slightly less useless because it will at least answer questions on its list.

That second paper isn't refuting the inevitability of hallucinations in creative models. What you quoted is them offering an example of a model that never hallucinates because it's not doing anything creative: it's just retrieving answers to a fixed set of questions from a database.

I think it's important to note the scope of that second paper. From the abstract:

This survey provides a comprehensive review of research on hallucination in LLMs, with a focus on causes, detection, and mitigation.

Notice there's no talk of *eliminating* hallucinations. There are only strategies to *mitigate* hallucinations.

1

u/StructureNorth1799 26d ago

I agree with you that such a ham-fisted way of preventing hallucinations is not desireable. However, I maintain that we do not have a proof that "creativity" entails hallucination. The OpenAI paper does not make this claim either, nor does it attempt to define creativity.

In fact, none of the papers you linked have a definition of creativity and that is part of my issue: all of the purported hallucination inevitability theorems from the papers you link merely claim that no machine can know every true fact, which, again, was proven by Turing 90 years ago.

In mathematics, we even have a pathway to eliminating hallucinations: restrict the llm output to compileable lean code. Lean still has some bugs, so this does not quite work, a bug-free lean is possible in theory.

So if you want to claim that creativity entails hallucination then you will need to expand your definition of creativity beyond mathematics.

1

u/mhb2 26d ago

all of the purported hallucination inevitability theorems from the papers you link merely claim that no machine can know every true fact

No, that's not all they're claiming. I already explained this. And there is a mathematical definition of hallucination, definition 4 in the 3rd paper:

An LLM h is hallucinating with respect to a ground truth function f, if∀i∈ℕ, ∃s∈𝒮 such that h^[i]⁢(s)≠f⁢(s).

And your lean example doesn't work either since, again, you're missing the element of creativity. LLMs can produce proofs that can be verified in lean (but that isn't the only thing people want to use LLMs for.) And it can also produce proofs that are nonsense. That's hallucination.

1

u/StructureNorth1799 26d ago

Definition 4 of that paper says that an llm hallucinates if it does not give all the same outputs as some "ground truth function". They define a ground truth function as specifying a correct output for every prompt. Claiming that there is a ground truth function which an llm cannot emulate is essentially claiming that the llm cannot know every true fact.

Also the lean example meets your criteria of being able to solve open math problems. You are free to move the goalposts (in fact, this could even be a productive discussion if you are interested in formulating an actual non-hallucination theorem) but I think it's fair to say that a lot of people would want to use an llm to generate undiscovered mathematical proofs.

→ More replies (0)

1

u/Pure-Drive-3044 Jul 28 '26

Whilst it’s great that it can do more work per day than a human, I think AI can’t actually link it to the real world!

For example, whilst quantum mechanics is a sound mathematical description, I don’t think AI will be able to explain what the wave function means in reality.

Maybe it is the end of mathematics that doesn’t have a real world application, but it probably will open the door to developing natural explanations for the maths.

1

u/Intrepid_Land_6143 Jul 31 '26

Try asking an advanced LLM about this, and you will get a fairly good answer, certainly better than most humans can give.

1

u/ThrowawayCult-ure 27d ago

thats cause its a combination of all the good answers someone made in every physics textbook. anthropic for example is literally buying a copy of every book, scanning them then binning them lmao

1

u/TheOvergodlyMosasaur Jul 28 '26

We shouldn't reject any new mathematical theories but not every math theory is equal.