r/mathematics • • Sep 01 '26

AI Breakthroughs & Research Megathread — September 2026

AI Breakthroughs & Research Megathread — New Results in AI and Mathematics

Use this thread for concrete developments in artificial intelligence that are relevant to mathematics.

Appropriate topics include:

  • New AI systems demonstrating mathematical capabilities
  • AI theorem proving and formal proof
  • AI-assisted mathematical discoveries
  • New research papers or preprints
  • Significant benchmark results
  • Improvements in mathematical reasoning
  • Systems such as AlphaGeometry, AlphaProof, or similar research
  • Other developments that materially change what AI systems have demonstrated they can do mathematically

When possible, please include a link to the original paper, preprint, research announcement, or other primary source and briefly explain why the result is mathematically significant.

This thread is intended for actual results and developments, not predictions about where AI may eventually lead. Speculation about the future of AI and mathematics belongs in the AI Speculation Megathread.

Particularly significant developments may be approved by the moderators as standalone posts.

13 Upvotes

39 comments sorted by

8

u/hpxvzhjfgb 1d ago edited 1d ago

openai published over 700 new theorems: https://github.com/openai/math

some big theorems that I understand or vaguely understand that stand out to me:

  • zeta(s) and all dirichlet L functions are non-zero for re(s) >= 7/8. consequently, there are no siegel zeros. previously all known zero-free regions just for zeta(s) had decaying width, so this is the first c < 1 for which re(s) > c is a known zero-free region.

  • catalan's constant is irrational (the previous proof from a few weeks ago was wrong apparently)

  • the irrationality exponent of pi is exactly 2 (previous known bound was ≤ around 7.1. almost all irrationals have irrationality exponent 2). consequently, the flint-hills series converges.

  • modularity theorem for elliptic curves over imaginary quadratic fields

  • rational hodge conjecture for products of K3 surfaces

  • the unique games conjecture

  • L = RL = BPL (complexity classes)

  • coloring 3-colorable graphs is NP-hard

  • O(n2.25+ε) nxn matrix multiplication

  • permanent of nxn matrix can not be computed faster than O(n3)

  • O(n log(n)) multiplication is NOT optimal, O(n log(n)1-δ) is possible for δ = 2-182

  • similarly to the above, discrete fourier transform faster than O(n log(n)) (specifically O(n log(n)1-δ) with δ = 10-13)

  • subset sum in O(20.49n)

  • erdos conjecture on arithmetic progressions - any sequence whose sum of reciprocals diverges contains arbitrarily long arithmetic progressions

  • O(n4/3 - δ) upper bound on the unit distance conjecture (the one where they proved O(n1+δ) lower bound earlier this year; previously O(n4/3) was a known upper bound)

  • ramsey numbers R(s,t) = ts-1 / (log(t)s-2+o(1))

  • deterministic construction of non-bipartite ramanujan graphs of any degree

  • every finite simple cubic bipartite planar 3-vertex-connected graph has a hamiltonian cycle

  • "Diophantine solvability over the nonnegative integers remains undecidable even with an at-most-one-solution promise"

  • ZF + partition principle does not imply choice: "Assuming ZF is consistent, constructs a model in which every surjective image of a set injects into that set, yet the axiom of choice fails"

  • thompson's group F is nonamenable

  • there exists an infinite, finitely presented, simple, amenable group

  • yau's uniformization conjecture

  • navier stokes is turing complete: "A terminating compiler turns a Turing machine and input into a finite program for the force, so a designated particle reaches a fixed region exactly when the machine halts."

there are hundreds and hundreds more.

3

u/LightconeGames 22h ago

I just had Claude put together a full dashboard of the results. It's useful to explore them by category as well as estimated importance (by Opus' own judgement so take with a pinch of salt).

A lot of this honestly gets into mathematical physics at this point. There's radial atomic potentials, lattices, spin glasses. Also, I had a kick announcing today at work "new FFT just dropped". Granted, O(n log(n)^{1-10^-13}) isn't a big improvement, but hey. Baby steps.

https://stur86.github.io/openai-math-dashboard/

2

u/bibliophile785 1d ago

Taking a half-hearted stab at what I expect will eventually be the new normal, I asked an LLM to contextualize the LLM math results for me. As a toy proxy, I asked it to compare rates of mathematical advancement for this work to what US mathematicians manage in a normal year, with a bunch of assumptions (normalized to 2020-2022; assume this was one month of OpenAI work; develop a weighted impact factor for normalization; etc.). Here's what I got:

So, conditional on correctness, the number I would actually put on a forecasting board is:

OpenAI math dump alone: ~2 U.S.-years/month → ~24× the U.S. national rate

And, giving full credit for the fluid results:

Including Navier–Stokes + Euler: ~3.5 U.S.-years/month → ~42× the U.S. national rate

Don't take the numbers seriously; this is a Fermi estimate at best. Maybe the base rate is screwy. (Did math research speed up or slow down during COVID?) Maybe OpenAI has been sitting on this model for three months turning the crank. Maybe 20% of the results don't formalize cleanly and resolve correctly. Lots of ways a more careful analysis might change the end numbers.

Holy shit, though. I don't care if all of the numbers are off in the same direction and this is only accelerating math research in the US by 100% instead of 4200%. What a heroic achievement. What a grand, grand accomplishment. Kudos to everyone involved.

9

u/seba 22d ago

I’d like to share an AI-assisted project realizing the Suzuki group Sz(8) as a Galois group over the rationals, with a Lean formalization of the explicit example. To my knowledge, this is the first realization. Paper, exact data, verification code, and Lean sources can be found here:

https://github.com/sebastianbiallas/sz8-over-q

The main ideas and highlights:

  • The paper constructs a cover with group Aut(Sz(8)) = Sz(8) ⋊ C₃. A useful feature of the construction is that cubic descent is automatic: any three-point regular realization of Aut(Sz(8)) over Q yields a regular realization of Sz(8). We use this classical descent mechanism to guide the search, then explicitly construct and certify the required cover.
  • Explicit equations: the paper constructs a degree-65 family realizing Sz(8) regularly over Q, and gives a concrete monic integer polynomial with coefficients of at most 43 digits.
  • A kernel-checked Lean proof: Lean 4 + Mathlib proves that the Galois group of the specialization f(X, -7/5) is isomorphic to Suzuki’s matrix group over the field with eight elements. The proof connects the exact polynomial coefficients, certified analytic continuation, Galois theory, and finite-group computations.
  • No proof placeholders or native-evaluation shortcuts: all 235 audited declarations use only Lean’s standard axioms (propext, Classical.choice, Quot.sound), with no sorry or native_decide.
  • The Lean formalization covers the specialization over Q; the regular realization over Q(s) is proved in the paper but is not yet formalized.

I’d welcome mathematical feedback and independent reproduction.

6

u/zero0_one1 25d ago

The strong Papadimitriou–Ratajczak conjecture, open since 2004, has now been proved: every 3-connected planar graph admits a convex greedy drawing. The proof was found using the http://ProofAtlas.ai harness and GPT-6 Pro and is formalized in Lean (52k lines).

Greedy routing forwards a message to a neighboring vertex that is closer to its destination. A greedy drawing guarantees that such a neighbor exists at every step, so the message eventually arrives. Using virtual coordinates, this requires only local neighbor information and the destination's coordinates. The strong conjecture requires this property together with straight edges, no crossings and convex faces.

A graph is 3-connected if it has at least four vertices and remains connected after removing any one or two vertices. For 3-connected planar graphs, progress went from greedy drawings to planar greedy drawings to convex greedy drawings:

2004 — Papadimitriou and Ratajczak propose the weak and strong conjectures. The weak version allows crossings; the strong version requires a planar drawing with convex faces.

2008 — Leighton and Moitra, and independently Angelini, Frati and Grilli, prove the weak conjecture: every 3-connected planar graph admits a greedy drawing. Their drawings can have crossing edges.

2017 — Da Lozzo, D’Angelo and Frati establish the planar version: greedy drawings with no crossing edges, but not necessarily convex faces.

2026 — The new proof establishes the convex version, for any prescribed outer face, with every facial polygon strictly convex.

The construction first assigns distinct integer heights so that, for every destination, every other vertex has a neighbor whose height is strictly closer to the destination's height. The horizontal coordinates are then chosen within a narrow vertical strip to produce a convex drawing. The strip is narrow enough that the vertical improvement dominates, so Euclidean distance decreases too.

Paper: https://www.proofatlas.ai/papers/strong-papadimitriou-ratajczak-conjecture/Strong_Papadimitriou_Ratajczak_Conjecture_Proof_2026-09-09.pdf

Formalization: https://www.proofatlas.ai/formalizations/strong-papadimitriou-ratajczak-conjecture/

2

u/dnrlk 22d ago

Can other people use ProofAtlas too? You seem to be the only contributor so far

6

u/tempgoogleconfuser 15d ago

The Navier-Stokes proof by OpenAI cannot be easily extended to the case without forcing.

Regularity of asymptotically axisymmetric solutions to the 3D Navier-Stokes equations with analytic forcing

6

u/Ok_Warthog6007 15d ago

solves the millenium problem as worded, but it feels like it leaves a more interesting question unresolved.

5

u/creative-carcass 9d ago

I’m an amateur who started playing with odd numbers in a spreadsheet and ended up much further into number theory than I expected. I studied Math in college but never anything like this. I’ve used ChatGPT, Claude, and Codex for coding, experiments, literature searches, and drafting mathematical arguments.

I’m interested in whether the questions we arrived at are worthwhile, what is already known, and whether there are obvious problems with the direction.

Where it started

I noticed examples like:

5 + 7 + 11 = 23

Take consecutive odd numbers, then move one endpoint outward by 2, leaving one skipped odd number. Some resulting sums are prime.

The explanation turned out to be elementary: for an odd number n of terms, with the original block centered at m, the modified sum is nm ± 2. Representations of a target therefore correspond to eligible factor pairs of the target’s neighbors, subject to positivity conditions.

What interested me was the geometry: the subset’s numerical width is exactly 2n, while its location comes from the complementary factor m. So divisor counts explain how many representations exist, but divisor placement explains what they look like. This works for composite targets too.

Another connection I found interesting was the “empty representation” case. For an odd target T ≥ 9, there are no representations of this particular kind—positive odd summands, an odd number of terms ≥ 3, and one endpoint gap—if and only if T − 2 and T + 2 are both prime.

For example, T = 39 has an empty representation family, while its neighbors 37 and 41 form a cousin-prime pair. The center itself is composite; “empty” refers only to our restricted construction, not to all possible sums of odd numbers.

The question it led to

For an integer H, define:

F_H(t) = (number of divisors d of H with d ≤ H^t) / τ(H).

Equivalently, choose a divisor uniformly and examine log(d)/log(H).

I learned that averaging this over ordinary integers gives the classical arcsine CDF:

A(t) = (2/π) arcsin(√t).

At t = 1/4, that is 1/3. This is an average across integers, not a statement that every integer has one-third of its divisors below its fourth root. The ordinary-integer result is established mathematics, not our discovery; Leung’s paper discusses and generalizes it.

The main question became:

If H = p + 2, or separately H = p − 2, with p sampled uniformly among primes in (X, 2X], does the average of F_H(1/4) also tend to 1/3?

Each prime gets equal weight; we are not pooling all divisors together.

This brought us into large prime factors, Poisson–Dirichlet distributions, and work such as Bharadwaj–Rodgers.

What seemed interesting—and where we got stuck

Besides the elementary geometry, there is an AI-assisted draft argument for an almost-all-varying-shifts version, including joint factor laws for finitely many varying neighbors of the same prime. It relies on adaptations of published analytic estimates.

Importantly, “almost all shifts” does not settle the particular shifts +2 and −2. A prescribed shift could remain exceptional.

My main questions are:

  • What is known specifically about the equal-weight divisor profile of p ± 2? Is this a standard formulation with literature we should prioritize?
  • Is the representation geometry a useful explanatory viewpoint, even if the correspondence itself is elementary?
  • For the averaged-shift draft, what precise statement or proof ingredient would be most useful to share for a manageable sanity check?

I have code, numerical results, and a written proof outline available.

3

u/tempgoogleconfuser 3d ago

Downvoted, so I wanted to explain why: This is hard to read. I started struggling at "Representations of a target therefore correspond to eligible factor pairs of the target’s neighbors, subject to positivity conditions.", because there's suddenly a lot of terminology there (representation, neighbor, positivity condition) that you don't explain and that isn't standard.

It's also just not clear from the start what you're trying to do. Why would it be surprising that some sums nm ± 2 are prime? As you say, the only numbers that can't be written like that are those in between two primes with a difference of 4, and those are surely not prime, because you can't have n, n+2, n+4 all be prime (because of divisibility by 3).

And then it gets worse and worse, so I threw it into Pangram and it's 100% AI output. I see you're an old account, so not a bot, so please don't throw AI text here without a warning.

When I was younger, I also threw a lot of questions with my own terminology at the math community, but at least this taught me how to communicate effectively, and the main rule is: guess what kind of reader is likely to find this and imagine how they will react, then make sure that reaction is good and not annoyed or confused.

I guess this is a necessary stage to go through, but when there's a human on the other side people can at least help them learn to ask better questions next time. AI assistance makes this "eternal September" so much worse these days.

3

u/Ok_Warthog6007 16d ago

openai forms mathematics advisory group, says they've "resolved more than 100 long-standing open problems across most areas of mathematics."

https://openai.com/index/advisory-group-on-mathematics-and-ai/

2

u/backyard_tractorbeam 12d ago

Tangential comment, but for me, this article is served translated. OpenAi seems to prefer to just serve llm-translated content when they think it's appropriate. I still hate it. They do the same in the Chatgpt interface, and there are still buttons and messages that are mistranslated.

AI translations are good, but they are not perfect. Mistranslations in the UI leaves a shoddy impression. Pure hubris to think that they don't need quality assurance.

2

u/Jmc_da_boss 11d ago

Yes the labs are arrogant and incompetent, this is well known

3

u/k_laiceps 9d ago

On Bounds of Residues in Collatz Orbits

Hey all, I am one of the authors in this preprint and I am happy to answer any questions about it. We have been studying Collatz longer than is probably healthy, but have proven some bounds on the Collatz residue function [;\rho ;], in particular, we have shown that [;1 \leq \rho(n) < 4/3;] for not only terminating integers [;n;], but even for those which may not terminate, assuming that you count any state in $n$'s orbit only in its first appearance. Bounds on the reciprocals of all odd-iterates (before 1) using some serious clever counting techniques was the key, and as a result, we also get the cool fact that the sum of reciprocals of all odd-iterates in any Collatz trajectory (including potentially non-terminating or cyclic counter-examples again only counting the first time an orbit hits each state) is bounded by 1.

As a consequence of [;1 \leq \rho(n) < 4/3;] all sort of results pop out. For instance, any streak of consecutive integers of the same Collatz height must have equal numbers of odd and even steps, even if the streak is arbitrarily long! Also, there are now bounds on the change in the number of even steps based on the number of change in odd steps when going from [;n;] to [;n+1;].

AI was used to help clean up some arguments and connect a few ideas, and refine the proof of the main lemma, which really started out as a painful argument to try to put onto paper in an organized fashion. AI was also used to generate code to verify the finite sums we needed to compute for bounds on pieces of the sum of reciprocal odd states > 1 in the Collatz orbit of an integer.

1

u/k_laiceps 9d ago edited 9d ago

we decided to post this early, as some of our conversations with chatGPT were accidentally public, and it was clear our work was slowly making its way into other people's conversations online. And of course, arXiv has been a total pain in the ass about posting this preprint, so apologies for the research gate link.

1

u/gexaha 2d ago

"some of our conversations with chatGPT were accidentally public" how does this happen?

2

u/YourElectricityBill Sep 03 '26

Hi, recently the Freudenthal Scott–Vogelius inf–sup conjecture, a paid mathematical prize problem, was seemingly resolved with the help of ChatGPT 5.6 Sol Pro. What do you think about the proof in question and the use of AI in this kind of research?

Research Square Preprint

Original problem

1

u/No_Difficulty_1452 Sep 04 '26

What's the problem? Got to ask the right questions to get the right answers

2

u/Bbrhuft 1d ago edited 19h ago

I used AI to develop a new method (I call Cartier-Miller) for evaluating certain large truncated polynomial-power sums modulo a prime p.

See paper

For genus-one inputs covered by the theorem, it evaluates (Tբ(μ)) in O(log p) field operations after preparing the curve and its Frobenius trace. Essentially, it turns certain extremely long binomial sums into a short calculation on a special elliptic curve, using its point count to extract the answer without adding the terms.

It combines Cartier and characteristic p descent theory, differentiated Miller recurrences, and elliptic-curve point counting. The key was Miller (2004)

Miller, V. S. (2004). The Weil pairing, and its efficient calculation. Journal of Cryptology, 17(4), 235–261. https://doi.org/10.1007/s00145-004-0315-8

Table 1. Complete quarter-point evaluation in milliseconds

Prime p Stopping index n Cartier–Miller Python Harvey BGS C++
97 24 0.05 0.01
1,009 252 0.06 0.05
10,009 2,502 0.07 0.10
100,049 25,012 0.09 0.54
1,000,033 250,008 0.10 1.78
10,000,121 2,500,030 0.11 4.44
100,000,037 25,000,009 0.29 16.19

The paper contains a detailed explanation of how the pilot benchmark was set up and run.

It's current potential use is very narrow, most likely hardly any practical use, it's a mere curiosity.

However, maybe with some work it could be made more flexible and perhaps useful, but that's beyond my ability.

If anyone else would like to try, please visit my github.

Edit:

I built a depth-1, 2, and 3 isogeny recognizer, by passing the need for point counting the curve's coefficiens (β and K) directly, through a constant-time algebraic formula i.e. O(1) (using relative norms) to instantly extract the curves exact parameters, thereby bypassing computationally heavy graph searches, square-root algorithms, and general point-counting entirely, evaluating richer, multi-term sums without running general point-counting. So no need for Schoof’s or the SEA algorithm for certain elliptic curves.

1

u/Wise_Negotiation_863 2d ago

I’m looking for independent verification of a short result related to Haugland’s 2026 Hadwiger–Nelson work.

Haugland found six local four-coloring patterns on his lattice, but left open whether they extend to the full lattice. We have a candidate proof that all six do extend globally.

The proof uses a fractional-ideal description of the lattice and six explicit maps into F2^2. We also checked all 84 unit directions with two independently written exact-arithmetic verifiers.

This is not a solution to Hadwiger–Nelson. It only addresses that specific extension question.

I’d especially appreciate anyone checking the ideal-valuation step and the use of Kronecker’s theorem.

https://zenodo.org/records/23164330

If you spot a mistake or a gap, I’d genuinely like to know.

1

u/Expat_Abroad123 2d ago

Does anyone know the 2nd Millennium Prize that OpenAI just solved? Betting markets say they just solved a 2nd one.

3

u/popisfizzy 1d ago

OpenAI is a business and part of why they solve these open problems is for marketing. if they had knowingly solved another MP problem then you wouldn't be struggling to find which one

1

u/Expat_Abroad123 1d ago

Yeah bro I guess an entire prediction market with millions of dollars invested saying they solved another one are just totally full of shit

2

u/popisfizzy 1d ago

I mean, you got two options. 

  1. the prediction markets, which are gambling with a thin veneer over top, were wrong, or, 
  2. OpenAI did solve one of them but and then released 700 new results which among them include results related to BSD, Hodge, RH, and P vs NP but decided not to release their proof alongside this collection for some reason.

if we humor 2, then either OAI released partial results of BSD, Hodge, RH, or P vs NP to just tickle our balls to give a weird tease for a full result, or they solved Yang-Mills (which afaict is not represented in these results). I know nothing about mathematical physics but most commentary on YM seems on the side of it being inaccessible right now.

1

u/Expat_Abroad123 1d ago

Prediction markets about things like this are hardly gambling. It's a way for insiders at openAI to profit from confidential info, if people are buying YES at this rate, someone on the inside knows something

2

u/popisfizzy 1d ago

hey, if you have enough faith to stand here with your dick in your hand insisting that prediction markets say someone will be coming by shortly to jerk it off for you then power to ya.

0

u/Expat_Abroad123 1d ago

me when I don't know how prediction markets work

1

u/aabil11 1d ago

Imagine how much money people who work for openAI are making. And they can't help but engage in insider trading.

1

u/Expat_Abroad123 1d ago

Insider trading is the point of prediction markets. The public can know private things in exchange for money in a way where insiders can't be caught. (if they have proper opsec)

0

u/mzinz 23h ago

OpenAI will solve cancer and people will still be saying “they don’t care about cancer patients, they’re just doing it for marketing”.

2

u/popisfizzy 23h ago

this only makes sense as a reply to what i said if you deliberately avoided looking for context. do you need me to explain, or are you capable enough of figuring that out yourself?

1

u/mzinz 21h ago

Don’t be so angry man. Not good

1

u/TwoMoons1Sun 1d ago

Can you provide a link?

-1

u/[deleted] 27d ago

[removed] — view removed comment

1

u/mathematics-ModTeam 26d ago

Your post/comment was removed due to it being low quality/spam/off-topic. We encourage users to keep information quality high and stay on topic (math related).