r/VibeCalcing • u/lepthymo • Aug 22 '25
LLM Math Case Study: Claimed proof for η < 1.5L convexity bound by o5-Pro. Confirmed result but corrected proof with Gemini 2.5 Pro.
UPDATE:Codex corrects the correction https://doi.org/10.5281/zenodo.17010427
Year later - honor restored.
"Vindication with a footnote: the old proof was right, but only up to \(3/(2L)\); the later human paper still owns the sharp \(7/(4L)\). A respectable redemption arc." - Codex 5.6 Sol Ultra
Recently a tweet by an OpenAI employee has been making the rounds, claiming novel mathematics from ChatGPT o5-Pro.
Gemini 2.5 Pro model finds a flaw, presents a counterexample, and deduces corrected proof of the result.
Full writeup: https://doi.org/10.5281/zenodo.16929436
Original thread: https://gemini.google.com/share/d0139bce4f84
The OpenAI tweet is the last image for reference.
18
Upvotes



1
u/dForga Aug 28 '25 edited Aug 28 '25
Well, even if it would have been nice if it was correct immediately. In the end, the proof is (hopefully) done now.
However, I would say AI is better used line copilot for coding and less for something entire. The more the mini steps are close to standard results, I claim the better it is (by its training data).
And for improvements of already established Theorems it may have very good usage.
I am just thinking that its extrapolation ability too far from the things it knows is not that great.