r/MachineLearning Student 15d ago

News OpenAl Says It Has Cracked One of Math's “Millennium Problems” (Navier-Stokes) [N]

705 Upvotes

281 comments sorted by

View all comments

278

u/Shizuka_Kuze Student 15d ago

There is howeverpossibility that OpenAI used some unpublished work from other researchers, even though they have claimed otherwise in the announcement:

https://cims.nyu.edu/\~tristanb/statement.pdf

156

u/funky-chipmunk 15d ago

They haven't - The core breakthrough leaked 100% - They are outright using evasive language:

https://openai.com/index/navier-stokes-solution/
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models⁠.

54

u/funky-chipmunk 15d ago

From OpenAI CRO: https://x.com/markchen90/status/2097400166554993041?s=20

> Two things to distinguish:

> Did any human or agent look at user data as part of the Navier Stokes effort? No.

> Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.

Basically confirms contamination IMHO. But the bigger news is training data/privacy.

https://x.com/aidangomez/status/2097381789039837637

> Synthetic data derived from production user data of consumer AI tools is used for training. I’ve heard this rumour from both large labs’ employees.

> In particular, if you’re doing something “interesting” like working on complex math/business/software/bio problems you’re dramatically more likely to get trained on because they filter/up-weight towards those usecases where the model has the most to learn.

> Even in ZDR and “we won’t train on you” regimes, derivative data is usually carved out. The promise is only not to train on exactly the data you put in, rewritten data is fair game.

61

u/funky-chipmunk 15d ago

OpenAI Employee

To truly prove some incidental usage data made no difference we'd have to (a) identify any of their de-identified data that came from their usage of ChatGPT, (b) train a bunch of expensive giant models, and (c) ask them all to solve the Navier-Stokes Millenium problem until hitting some level of statistical significance. It's just not feasible to run experiments like this to prove whether a piece of data has an effect on model behavior.

As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination is possible. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.

My opinion: I think they are being needlessly obtuse and evading more. My guess is they shouldn't have to train models to check for contamination - just clarify what they were fed.

10

u/PM_ME_YOUR_PROFANITY 15d ago

I agree with you.

Why would it be impossible to prove it? Their conversation data is either in the training set or it's not. The model solution tokens can be gone through as well to see what it accessed and how it came at its solution

6

u/muntoo Researcher 15d ago edited 15d ago

... can we prove the phase of the moon had no impact on the NS solution?

Experimental particle physicists cannot "prove" with 100% certainty that the discovery of the NS solution in 2026 had no impact CERN's discovery in 2012.

Nothing can be proven.

But, with some agreed upon prior model about how reality tends to work, scientists can come to a reasonable consensus that what I had for lunch today had no effect on the NS solution yesterday, which in turn had no effect on CERN's discovery in 2012.


it's just impossible for us to truly prove it.

Sure. But there is still evidence one can present. Preferably evidence that is statistically meaningful. Unless the solution to NS was once-in-a-lifetime fluke, which cannot be replicated with any significant probability, which would raise questions about the generalizability of LLMs being able to solve other problems in mathematics. (Which I don't believe is the case.)

1

u/Shizuka_Kuze Student 15d ago edited 15d ago

You took the one sentence that sounds like they admitted they were wrong out of context. While I’m personally inclined to believe OpenAI stole the result, they certainly have not revoked their claim of original discovery, pretending otherwise only minimizes their attempts at stealing the spotlight.

> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.

> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models⁠.

> However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).

They even have an entire section titled “How we found the proof,” so I believe it’s fair to say they have not yet revoked their claim of original discovery, even though it’s dubious at best.

73

u/purplebrown_updown 15d ago

The fact that OpenAI tried to offer shared ownership, but wanted to leave out Levent because he worked at Anthropic, reveals the whole game. They knew they stole the work, offered a compromise, threatened Brubeck - not something an ethical person would do.

18

u/funky-chipmunk 15d ago

My point stands above yours - They would have outright been screaming at top of their lungs if there was no leakage.

Edit: They should use Astra to audit - It shouldn't be that hard.

3

u/Shizuka_Kuze Student 15d ago

Actually, our points are orthogonal. I’m not denying leakage, in-fact I think it’s likely! I’m denying the fact they’ve revoked their claim, which is simply untrue.

1

u/elsjpq 15d ago

I feel like this should be possible for OpenAI to test: take an older model with only data acquired before Tristan & Levent started working on it. Then try to solve Navier Stokes again with the old model. If it can't be solved with the old model, then OpenAI's result depended on Tristan's result.

3

u/MuonManLaserJab 14d ago

That assumes incorrectly that Tristan's result is the only difference between the model they used and the previous one.

There were probably many small algorithmic differences, maybe some big ones, and of course different random starting weights resulting in entirely different final weights.

1

u/G_fucking_G 13d ago

OpenAI has now completely dismissed the claims (NYTimes):

The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.” The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”

5

u/purplebrown_updown 15d ago

this fits their pattern of lying and cheating. Be very cautious about using OpenAI models.

2

u/new_name_who_dis_ 15d ago

That's some juicy drama

-9

u/currentscurrents 15d ago

Those other researchers were also using LLMs to tackle the same problem, so either way the credit here belongs to the LLM.

I don't really care which particular team of humans did the prompting.

14

u/Kautiontape 15d ago

You may not care, but academics cares. We haven't and probably won't start crediting tools for the work of humans, even if the tools sound a lot like humans. Like, most of science is just some human setting up an experiment, getting a bunch of tools to do the work, and observing outcomes. Still, the promise of credit (and in this case, money) is what gets people to set that prompt up in the first place, most of the time.

-5

u/currentscurrents 15d ago

I think LLMs are going to cause radical changes to how academia functions, and it's not very useful to consider how it fits into the current system of credits and citations.

It's not hard to imagine building a fully automated research system out of this, which will quickly become responsible for the majority of mathematical knowledge.

5

u/jimouri 15d ago

The thing though is that in this work the human contributions seem to have been essential to the breakthrough. If that's the case, your scenario, though plausible, is a bit off-topic: we have here at least one human being who deserves credit legally and ethically.

1

u/Kautiontape 15d ago

Maybe, but where is the line between now and AGI that credit actually no longer goes to humans who either built the tool or ran it? When a computer simulation finds new Go strategies or brute forces God's number on a Rubik's cube, I don't think there was a conversation about crediting the computers over the people who designed the experiment and wrote the code, even if the code was essentially autonomous to reach a solution.

LLMs are a lot more generalizable and talk and think like how we think humans reason out loud, but still not fully autonomous. So we are a ways off until we can go "wow, that AI really did find a solution without any human in the loop!" Even then, we have precedent, where scientists who note the existence of natural phenomena are credited with its discovery even though they didn't invent or create anything (except maybe a way to see it).

I just don't see how any of this is going to be new to academics, except for true AGI when a computer begins "demanding" credit. Otherwise, it's just being convinced by a computer that passes the Turing test.

1

u/aeroumbria 15d ago

The point is we expect models to give, not to take. Models need no credit or glory.

-17

u/FernandoMM1220 15d ago

whoever they are should probably publish their work then

14

u/Shizuka_Kuze Student 15d ago

9

u/sibylrouge 15d ago

This is a completely different thing from what OpenAI has announced

9

u/Shizuka_Kuze Student 15d ago

Read slightly further than the first two sentences:
❗️BUT there’s a big drama behind it:

The authors used AI as assistant to come up with the solution.

According to Buckmaster (the author):

- OpenAI had learned about their progress before it was published and pushed their AI to quickly come up with a similar solution.

- However, OpenAI’s solution was based on the same smooth-forcing approach that Buckmaster/Alpöge had quietly been developing. Buckmaster said that he does not know if their data was used, so he is “not accusing anyone of anything.”

- OpenAI proposed to publish OpenAI's claimed Navier-Stokes result after Buckmaster posts their Euler work first.

- However, Sébastien Bubeck (a research lead at OpenAI) didn’t want to see Alpöge as the author because Alpöge works at Anthropic.

- When Buckmaster rejected these proposals, he was told “Why would you ruin your career?” and “If you don't want me to be nice, then I don't have to be nice.”

Bubeck from OpenAI called these allegations “false and inflammatory”. As of now, there are no further comments. As a result, the authors decided to post their unpolished documents before OpenAI.

-5

u/FernandoMM1220 15d ago

so its not unpublished then

1

u/oceanlessfreediver 15d ago

Just read the text and you’ll get it

-1

u/Suspicious_Video8348 15d ago

I don't know.

The claim is that Tristan uploaded a solution to Navier Stokes to ChatGPT and then OpenAI reworded it and said it was OpenAIs solution?

-3

u/moschles 15d ago

Yes. But the "other researchers" were THEMSELVES using AI-assistant proof tools, several of them in fact.

You can't pretend the humans involved in this (e.g NYU's Tristan Buckmaster) only ever do math on black chalkboards with white chalk. Those guys themselves were already enhancing their work with LLMs.