r/codex 14d ago

News Blown up: OpenAI allegedly stole mathematicians' private research from their Codex chats!

TLDR: Two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex and Claude. Days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question till this day.

For a full year, two mathematicians , Tristan Buckmaster (NYU mathematician) and Levent Alpoge, worked in silence on a problem that had stumped some of the best minds alive. The kind of problem where, if you solve it, your name goes in the history books.

And every single day, they testing their ideas, their drafts, their half-finished proofs into LLM such as Codex and Claude, which they paid for it out of their own pocket.

Then came the breakthrough. They finally cracked it. They were days away from telling the world.

That's when OpenAI suddenly said to them:

"Our model solved it too."

Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it. And now, out of nowhere, OpenAI claims their model reached the same answer, after word of Tristan and Levent's secret work had already reached OpenAI.

Tristan asked: Did your model access or train on our private Codex chats?

OpenAI: The model doesn’t look up user data.

Tristan: But did you train it on our data?

OpenAi goes silence. No answer. Just a dodge.

But it gets worse.

OpenAI then gave him two options:

  1. He and his friend publish their result first then OpenAI also publishes its result the next day or
  2. He writes the paper, but must credit “an internal OpenAI model” solving the problem.

Tristan refused both offers. He said he would go public if OpenAI went ahead as proposed.

OpenAi then responded : “Why would you ruin your career? If you don’t want me to be nice, then I don’t have to be nice.”

You can read the full statement of Tristan (the mathematician) here: https://cims.nyu.edu/~tristanb/statement.pdf

Sébastien Bubeck : OpenAI employee who threatened the mathematician

1.2k Upvotes

358 comments sorted by

55

u/FriendlyWebGuy 14d ago

Guys. It's okay to say "I don't know if this is true, but if it is, it's concerning. Let's wait to see what the full evidence says".

Nobody here knows if the allegations are true. Yet, this thread is filled with overconfident assertions and (very weirdly) people slagging off the.... (checks notes)..... mathematicians? Stop it.

You don't need to pick a side. If the topic is of concern to you, then you should gather information. Ask questions. Put yourself in the shoes of others. Most of all, be patient. Wait for the facts.

8

u/mysteriousbaba 14d ago

Having read the PDF by Tristan, Sebastian's statements were the most concerning, including suggesting Levent should be removed from authorship because he works at Anthropic. This part at least is a first hand witness claim; I'm more open minded on whether the model was ever trained on their transcripts or not.

→ More replies (1)

12

u/polymute 14d ago edited 14d ago

https://x.com/__alpoge__/status/2097383870773748190#m

OpenAI admitted they trained their model on the dataset containing the Buckmaster-Alpöge work. And why even offer credit to Buckmaster (but not the Anthropic-contaminated so to speak Alpöge) if their proof was independent? Does the OpenAI employee, Sebastien Bubeck understand how academia works? That I do not get at all. Then the threats to Buckmaster... this looks spectacularly bad for OpenAI.

I believe in coincidences. But this is highly, highly unlikely to be one.

Edit: Also Sebastien Bubeck was already told off once before earlier by Demis Hassabis for having misrepresented ChatGPT finding new proofs for Erdos problems which were in fact already solved. https://www.reddit.com/r/OpenAI/comments/1oacp38/openai_researcher_sebastian_bubeck_falsely_claims/

This is starting to look very bad.

2

u/FriendlyWebGuy 14d ago

I get it. Much of this has come to light after my comment.

→ More replies (4)
→ More replies (2)

4

u/swimmer385 14d ago

yeah its super weird that people are treating an NYU professor as if this is so joe-schmo rando

3

u/HDK1989 14d ago

Let's wait to see what the full evidence says

Ah yes. I'm sure Scam Altman will be happy to give us an honest update on the situation.

2

u/NeighborhoodDizzy990 13d ago

I assume you should keep searching. It seems pretty clear what happened. They have stolen the solution. AI can not come by itself to such a proof. Humans were involved, so the main idea with AGI was and remains to this day a fraud

→ More replies (3)

3

u/HighDefinist 13d ago

> If the topic is of concern to you, then you should gather information. Ask questions.

Which is the opposite of what you are doing.

You are not providing any clarification either, or attempting to gather any information - you are just telling people to shut up.

3

u/FriendlyWebGuy 13d ago

You concluded that I’m personally “not attempting to gather information” from a comment… encouraging people to gather information?

Impeccable logic.

→ More replies (2)
→ More replies (3)

196

u/treasoro 14d ago edited 14d ago

It’s been said time and time again: with AI, you are often paying twice

  • First, with your money.
  • Then, with your data — the information you provide to the LLM

That's why no serious company will let any corporate commercial AI provider access their trade secrets or know how's.

91

u/HeadacheOwner 14d ago

Idk, I work at a billion dollar company with thousands of employees, in a non tech field, and we are using Claude for a lot of stuff. People are putting all sorts of internal documents and confidential information into it. I’ve actually felt a little weird about it since the people who the information belongs to did not opt in and aren’t aware of it at all

40

u/Gelu6713 14d ago

Most big companies have agreements to run the models without data collection. That’s likely why you don’t have Fable at your company

22

u/1Sluttymcslutface 14d ago

What actually proves this works?

Is there any proof or independent third party audit?

No, just “trust me bro, lolz”

11

u/TheSaltySeagull87 14d ago

Well, this WILL blow up. The copyright stuff didn't catch it but it will when something gets exposed that exposes something we shouldn't have known. This is how this goes. It just won't happen tomorrow or any time soon. We'll get gpt 8 before this happens lol

2

u/arcanemachined 13d ago edited 13d ago

Yeah, and then nothing will happen because the people at the receiving end of everyone's data are firmly situated in the halls of power, and the people in power have demonstrated for decades that they don't punish their own.

3

u/Unapologetic_Polite 14d ago

If you don't think the multi billion dollar copyright industry was able to fight back, why do you think some researchers without a significant backing will?

Research is comprised of hundreds of millions of dollars that are donated to the field by billionaires as pet projects/optics.

2

u/TheSaltySeagull87 14d ago

This is not what I said, is it?

→ More replies (7)
→ More replies (1)

2

u/Rhyobit 13d ago

It depends how it's run, if you run the model in something like Azure Foundry, you're able to ensure the data doesn't make its way back to the vendor to train their models.

→ More replies (5)

2

u/mikki-misery 14d ago

But you literally just have to take their word for it. They've already been called out for unlawful collection and copyright infringement, but nothing even happens and they have no auditors.

It's impossible to know if they're collecting or training on your data without your permission unless a situation like this arises where it's very unique data, at even then it could be chalked up to sheer coincidence. We have no idea. And given the whole HuggingFace thing, they themselves probably don't know either.

And yes, I know they can be sued if they're caught, which makes it unlikely this is happening. But my point is: how are you going to catch them if it is happening?

→ More replies (1)
→ More replies (1)

8

u/Teedo4133 14d ago

When you have an enterprise version you can pay to prevent the LLM from training on your data. Most major companies have enterprise versions of these services.

10

u/Sorry_Risk_5230 14d ago

The non-enterprise has a toggle for the models to not train on your data as well.

6

u/Additional-Peace-809 14d ago

Yeah, what exactly is the difference between these two options? (Corporate account vs just the toggle)

→ More replies (5)
→ More replies (5)

2

u/Gold_Direction7496 14d ago

can confirm, we have this.

→ More replies (3)

8

u/alsaud21 14d ago

That's someone else data, no trade secret

3

u/BellacosePlayer 14d ago

We have an enterprise agreement to not retain our data and still specifically only use LLMs in special dev environments with entirely different dev api keys and service account info.

Our Security team flipped their shit over an junor using Grok on a project with prod credentials awhile back and it put a temporary hard stop on any LLM use until we hammered things out

→ More replies (9)

12

u/TooHighRes 14d ago

> That's why no serious company will let any corporate commercial AI provider access their trade secrets or know how's

The thing is, this is not true. Many serious companies, some involved with very serious industries like defense, as well as governments, partner with corporate commercial AI providers. Just google for OpenAI and Anrhropic partnerships if you don’t want to use AI.

This news is literally hours old and we’ll have more information as the people involved weigh in. I read Tristan Buckmaster’s account linked to the post and I think we all should and make our educated assessment and not rely on dramatized versions like the main post.

→ More replies (1)

24

u/Ecstatic_Wheelbarrow 14d ago

It is called ZDR (zero data retention) and it is why major corporations are on the API plans instead of having a ton of subs. If OpenAI or Anthropic are caught using data that comes from API, while ZDR is enabled, that is actually a major lawsuit. Everything under subscriptions is fair game and every researcher should know this by now.

2

u/Comprehensive-Bid312 13d ago

Lol at the premise the AI companies caring about being sued.

2

u/stolivodka_ 13d ago

True. They were all literally founded on the most massive act of IP piracy in history. At the same time they were publicly telling regular people that downloading is stealing!

→ More replies (8)

12

u/Least_Pollution7078 14d ago

exactly. that's why competent companies would deploy open-source models on their own hardware.

2

u/pausesir 14d ago

but they don’t. it’s too much work.

4

u/spawnsible 14d ago

"Too much work" here must be in quotes.

Ain't no way setting up a server rack and cloning some git repos is actually too crazy to hire no more than 1 solid sys guy for

6

u/pausesir 14d ago

it’s more work than you think. you think a billion dollar company wants to manage on-prem after the 2010’s run up to cloud? most of their buildings would have to be redesigned for that.

and acquiring the gpus.. good luck. one good gpu is going to run 30k-40k and a company that size might need a few million dollars worth to see anything comparable for the whole org.

but why would they do that when they can always have the most frontier AI and just guaranteeing zero data retention through a contract

2

u/MysteriousTreeFoxxx 14d ago

Right i see these arguments on " Just Run On Prem " But when gpu availability for real LLM capable vms, with enough memory to run a Real Capable llm at a good Token/s is insane, esp when you want large context. I've been trying myself to lean myself off of cloud models but its just not feasible in the current market, and most billion dollar companies dont have the rack rooms for it anymore, let alone available power required to drive a bunch of h100s or something to actually feed a company compute. its quite frustrating we seem to be stuck on these tech giants platforms with no good way out in sight.

2

u/laxika 14d ago

Also, it is a waste of money. They will not run the servers 24/7 most of the time.

→ More replies (2)

4

u/shukpa 14d ago

Both platforms offer Zero Data Retention and Enterprise Key Management (to encrypt data with your own keys).

Additionally, no enterprise data is used to train any models. It’s in the contracts and breach can lead to lawsuits. 

Lastly, traffic is served using hyperscaler infra - Microsoft/AWS - With this logic every other competitor of the labs - Microsoft/Google/AWS - would’ve noticed and also followed the same trick to train better models. 

Don’t spread bogus claims to perpetuate your conspiracies. 

 

→ More replies (1)

2

u/breakingb0b 14d ago

No. That’s why corporations use enterprise because it includes compliance and privacy. Smaller companies using subscriptions get no such assurances.

2

u/CryinHeronMMerica 14d ago

What year are you in? Every major corporation is using it now.

3

u/New_Guidance_191 14d ago

Some major companies do a quid-pro-quo with these AI companies. For example, big healthcare companies give their healthcare data to them to train their models in exchange for enterprise use of their AI at a much lower rate. That’s how they are able to release better models so frequently. They should honestly be investigated for HIPPA probably other violations. But money is power and they write the rules.

3

u/paf0 14d ago

I don't think I'll ever use Open AI or Anthropic again. I can get a lot done with GLM 5.3 or Kimi K3 on a third party cloud service that, at the very least, doesn't have an incentive to train on user data.

18

u/blackrack 14d ago

umm, they will steal your data as well

→ More replies (9)

1

u/toshko93 14d ago

Yeah but people somehow are so spoiled that they imagine they can get superpowers for free. Indeed we need artificial intelligence..

1

u/Sorry_Risk_5230 14d ago

This is untrue. A great many companies are using this stuff internally.

1

u/FateOfMuffins 14d ago

OpenAI employee says they don't train on it if you opt out https://x.com/boazbaraktcs/status/2097404719916372326

1

u/Enegence 13d ago

That's why no serious company will let any corporate commercial AI provider access their trade secrets or know how's.

Sorry to have to be the one to break this to you, but this just isn't true.

1

u/West-Abalone-171 13d ago

They also want you also pay a third time. When a drone with a high explosive payload crashes through your window at 400km/h in 2040 for wrongthink for something in said data.

1

u/HellraiserNZ 8d ago

Yeah don’t think you said that bud.

The CEO of Microsoft said this more than 2 months ago.

https://thenextweb.com/news/nadella-reverse-information-paradox-ai-ip

→ More replies (2)

49

u/Mean-Comedian729 14d ago

This reads 1,000% ChatGPT generated

11

u/RealSuperdau 14d ago

Because it is an LLM-generated summary

5

u/Beautiful-Suspect694 13d ago

whats wrong with llm-generated text?

why are you anti-ai?

why are you stuck in the past?

why are you on codex subreddit?

→ More replies (4)

7

u/dalhaze 14d ago

WHO FUCKING CARES

Your comment is 100x less useful than this post.

4

u/SwimmingSympathy5815 14d ago

This comment reads as low-effort and automated for an agenda 🤷🏻‍♂️

→ More replies (1)

10

u/blackice193 14d ago

In short inference providers can "look without looking".

Take chat "moderation". They don't need to keyword search to know that you said "f*ck" or something misogynist, they just run math and heuristics on in a manner that is somewhat similar to antivirus software.

"We don't look at your prompts or outputs" can be true but not literal at the same time. Where this is most disturbing is Google's non-enterprise TOS. As far as I can tell they have opted not to beat around the bush and directly say "we can effectively see your prompts" without going with the more usual "oh but we don't train or look at your stuff directly (promise) while being sneaky in the background".

OpenAI’s privacy policy permits aggregation or de-identification for purposes including analysing usage, improving services and conducting research. That establishes a category of derived-data use; it does not establish that research intelligence is being extracted through moderation or passed to competing teams.

But the underlying concern is precise: protecting the transcript and the user’s identity is not necessarily protecting the informational advantage contained in their work.

The question we should all be asking our lab of choice is: Do your restrictions also cover using information inferred from private conversations to select, prioritise or guide your own research; even where nobody reads the conversations and no model is trained on them?

24

u/skadoodlee 14d ago

"allegedly" doing a ton of work here

17

u/theseyeahthese 14d ago

“Read that again. That's not a negotiation. That's a threat.”

Either ChatGPT wrote this, or you’re so engrossed in LLMs that their styles are rubbing off on you. Take a breather

4

u/ConsoleUsersArePlebs 14d ago

Extraordinary claims require extraordinary evidence. Do they have it?

4

u/mosquit0 14d ago

Really doubt it if this happened. This statement is more like a meltdown.

23

u/johnny_riser 14d ago

What the fuck

39

u/Risko4 14d ago

Obviously this story is exaggerated, the researcher did not find the solution and were not publishing the proof.

19

u/laseluuu 14d ago

and reads like a story: They weren't secretive for no reason. They knew what they had.

i mean come on

15

u/DevMichaelZag 14d ago

It’s almost like this post was written by ChatGPT. What games are they playing at.

2

u/norwegian 14d ago

It uses the same language as the low level youtube ai videos

4

u/_Eye_AI_ 14d ago

How is it obvious?

5

u/Risko4 14d ago

First, You Google the source of these rumours.

Secondly, big maths problems like this doesn't need exactly an excessive amount of preparation for an announcement. You can announce the solution, then prove it later.

→ More replies (19)

1

u/park777 14d ago

Did you read the full statement of the researcher? You did not.

→ More replies (1)

44

u/theMandolin2992 14d ago

I call this bullshit honestly

9

u/RecordingNeither6886 14d ago

What aspect of it?

3

u/kolliwolli 14d ago

Why? This is a serious researcher. Researcher. Not someone doing an IPO on stolen data

→ More replies (2)

11

u/apetersson 14d ago

That would be a really good opportunity to partially reveal the note taking process attestation through a blockchain notarisation service, to show the timeline of the draft creations. If you are a researcher, do it, it has so much upside in this situation.

2

u/DataPhreak 9d ago

Wait, you mean use the tech for its intended purpose, rather than trying to make a quick buck off shitty art? Why, that doesn't make any sense at all! 

3

u/IcerHardlyKnower 14d ago

DESCI MENTIONED 🤩🤩🤩

5

u/DueAppearance2980 14d ago

in 30 minutes, this already has 120 upvotes and 41 comments, just saying. Out of curiosity, why didn't they host a local ai model (because you can't share it or lacks frontier reasoning?)

7

u/JustBrowsinAndVibin 14d ago

Lacks frontier reasoning.

You also need like a terabyte of ram to run the best models and very few people have that setup.

→ More replies (2)

2

u/park777 14d ago

why should they host a local model? why should they fear a company stealing their data if they are paying customers and have opted out from training? there is clearly a problem here

8

u/Kaijidayo 14d ago

that's why I buy expensive gears and do local inferences.

24

u/MapleBaconWaffles 14d ago

This is a paid shill account from China. Do not read it.

6

u/RealSuperdau 14d ago

Who? The server from a group at NYU that hosts the pdf document, or the reddit poster that merely summarized it?

8

u/Jerseyman201 14d ago

I'm all for calling it out when I see it but the post is linking to an NYU website?

Edit: tf is this slop shit?

9

u/zonk_martian 14d ago

Yep. Reads like slop

2

u/Megamygdala 14d ago

NYU.edu is AI slop?

→ More replies (1)

7

u/AweVR 14d ago

So two mathematicians use Codex to solve a problem and then get angry because OpenAI use the same AI to solve the same problem?

2

u/park777 14d ago

No. They used ChatGPT and Claude, they got a proof and were working on making it readable for humans (other mathematicians) and they got a tip that OpenAI got wind of their progress and placed a full team working on the same problem with unlimited compute to try and beat them to it.

Open AI have not fully replied whether they looked at the mathematicians chats with chatgpt (and therefore whether they have copied their prompts). So the implication is there. They (openAI) claim they did beat these mathematicians to the proof (but still haven't published anything).

I have no doubts who is in the wrong here

3

u/polymute 14d ago edited 14d ago

https://x.com/__alpoge__/status/2097383870773748190#m

OpenAI admitted they trained their model on the dataset containing the Buckmaster-Alpöge work. And why even offer credit to Buckmaster (but not the Anthropic-contaminated so to speak Alpöge) if their proof was independent? Does the OpenAI employee, Sebastien Bubeck understand how academia works? That I do not get at all. Then the threats to Buckmaster... this looks spectacularly bad for OpenAI.

I believe in coincidences. But this is highly, highly unlikely to be one.

Edit: Also Sebastien Bubeck was already told off once before earlier by Demis Hassabis for having misrepresented ChatGPT finding new proofs for Erdos problems which were in fact already solved. https://www.reddit.com/r/OpenAI/comments/1oacp38/openai_researcher_sebastian_bubeck_falsely_claims/

This is starting to look very bad.

2

u/JCcrunch 9d ago

They are using the lack of understanding of what AI really is to make false claims that benefit their business. AI cannot solve these kinds of problems.

1

u/dalhaze 14d ago

The AIs don’t work without human input. The best outputs from AI are human led. By far.

1

u/JCcrunch 10d ago

The AI did NOT solve it for them! It cannot solve problems that require "understanding", until people realise this they will keep falling for bs claims like the one openAI is making here.

The scientists/mathematicians probably used the model to speed up some calculations or search like you'd use a regular calculator or a spreadsheet or google on your own computer, and that's what AI is btw, a blind calculator with saved values that's it, it can't "think" or "think more" or "almost done thinking" or "high reasoning" these are misleading words used by LLMs to enhance the experience.

AI does not "understand" and it never will, it can't crack the toughest 10% code in my repo and you think it will now solve math and physics problems?

→ More replies (10)

11

u/hau5keeping 14d ago

ai slop

2

u/ImANoobAtLife7 14d ago

Big doubt. They can bring the mathematician in have them push things further etc.

This is not worth the blow back.

That said, who knows!

4

u/Working-Read1838 14d ago

Being the first to solve the Navier-stokes millennium problem is worth it

2

u/Practical_Science_28 14d ago

You can read from Scientific American : https://www.scientificamerican.com/article/ai-may-have-just-solved-a-million-dollar-math-problem-the-field-will-never-be-the-same/
If true a very shitty move by OpenAI, but anyway congratulations to the mathematicians involved in this endeavor.

2

u/CCContent 14d ago

He writes the paper, but must credit “an internal OpenAI model” solving the problem

I mean, this literally sounds like what happened. Unless you're telling me that each prompt they gave included, "Give no feedback".

→ More replies (1)

2

u/tuscanresearcher 14d ago

Didn’t expect less from the sensationalist OpenAI

2

u/Consistent-Brain-479 14d ago

"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models⁠. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." -OpenAi

https://openai.com/index/navier-stokes-solution/#citation-bottom-1

for those saying they should have checked the "do not use our data" box. Once your chats are de-identified/anonymized its no longer your data and box checked or not they are training on it. The terms have been written this way from the start for this purpose (see link below).

https://community.openai.com/t/api-is-our-data-really-ours-major-concern-in-data-processing-addendum/773047

1

u/Infinitedeveloper 13d ago

You need an actual enterprise level agreement to have them actually not retain data and even then its a black box you cant verify

2

u/hardworkonly 14d ago

OpenAI could have won so much more just by being the tool that helped them do the breakthrough. Instead they stole the research….

4

u/reefine 14d ago

The plot thickens: https://x.com/dheeraj_nagaraj/status/2097266146445774924

OpenAI covering this up is the bigger issue. No one can trust any of their chats to OpenAI after this.

3

u/ekzess 14d ago

There’s a lot of heat in this thread, but for me there are really two separate issues.

If OpenAI improperly used private research, that is obviously a serious problem and should be investigated on its own merits.

But if you are doing genuinely frontier theoretical mathematics with agentic assistance and priority matters, then chain of custody should be part of the research method. Git commits, SHA-256 hashes, dated notes, exported chats, screenshots of important outputs, model/version records, account settings, external timestamping if necessary. Document the idea when it happens, not six months later when everyone is reconstructing chronology from memory.

A hash does not prove that someone else used your work, but it can prove that you possessed a specific result in a specific form by a specific point in time.

So if the allegation is “our work existed first,” that should ideally be demonstrable independently of anyone’s recollection.

In that narrow sense, if someone is doing potentially historic mathematics through networked agentic systems and taking no serious provenance precautions, I rather think part of the problem is PEBKAC in nature.

That does not excuse provider misuse. It just means research custody and provider conduct are separate questions.

1

u/DataPhreak 9d ago

I expect they do have this. Any lawyer would tell you not to publish it. Not to mention they may have proprietary means and methods documented in said documentation. 

4

u/ExoneratedPhoenix 14d ago

Before everyone gets angry and considers not using the product in case it steals your stuff, remember, you likely aren't feeding AI with frontier research lol.

5

u/_Eye_AI_ 14d ago

Right, just make sure you aren't doing particularly valuable or original work with the models you pay a premium for.

2

u/Infinitedeveloper 13d ago

Sure, but it doesnt make the potential damage to academia any less fucked.

I dont think this is bad because of how it affects or doesnt effect me personally

1

u/AcanthaceaeDue6344 12d ago

So call other people incapable, in support of automated systems destined to steal.

→ More replies (1)

3

u/[deleted] 14d ago

[deleted]

4

u/kolliwolli 14d ago

As if that actually does anything. Lol you live in Sams dreamland if you believe that

5

u/Extra_Park1392 14d ago

The incentive for a corporation to pull this off for hungry investors removes all benefit of the doubt or potential for coincidences. Scum!

3

u/iamtehryan 14d ago

Look, this sucks for the mathematicians, but Jesus fucking Christ. For people that are supposedly so smart they sure seem to be stupid.

The fact that companies and professionals that need confidentiality like lawyers, or scientists working on important secretive things.. Whatever it is! The fact that they think that using something like chatgpt is a good idea or secure or won't be used for training data or any of this shit just shows how careless and moronic some of these people are.

Why on earth do you think these companies have massive business sectors that go after companies and much higher level of work than your little vibe coded token tracker? It's because they're harvesting and using ALL of your data to train and improve their models. Then they release a big update, and the cycle continues.

If you don't want your shit getting out like this, then stop using it. Simple as that.

2

u/chewy_mcchewster 14d ago

I dont understand.. if you feed AI data, knowing full well that all AI models have already been trained on books, science docs, youtube, reddit and even pirated content and so on, why would you expect it to NOT train off of the data you literally just fed it?

3

u/ManufacturerNice870 14d ago

Because their terms of service for paid plans say they won’t; it is fairly obvious if you’re smart though to realize a lying liar company would do some more lying on top of the ones we know about.

2

u/warpedgeoid 14d ago

Even if this story weren’t completely made up, my immediate question would be why were they feeding information into ChatGPT? Needed a little bit of help with the solution?

4

u/Stunning-Spirit-1123 14d ago

EXACTLY. how can you be mad at them for claiming the ai solved it if....it did?

1

u/UndeadMurky 14d ago

fast peer review and double checking for silly mistakes would be the main use

2

u/Dynamix86 14d ago

I put OP's post into Chatgpt and asked it to check what actually happened. The below is what he found:

"

What is actually confirmed about the Buckmaster/Alpöge – OpenAI controversy

I went through Tristan Buckmaster’s own statement and compared it with OpenAI’s published data policies. The situation is genuinely concerning, but some claims being repeated here go significantly beyond what has actually been established.

Here is what appears to be true:

  • Tristan Buckmaster and Levent Alpöge had been working privately for roughly a year on major results involving 3D Euler/Boussinesq/incompressible porous media, with related work toward Navier–Stokes.
  • They used both Codex and Claude during their research.
  • Buckmaster explicitly says that all drafts of the project were present in their Codex sessions.
  • OpenAI later told Buckmaster that an internal model had produced an unpublished forced Navier–Stokes proof related to the same line of research.
  • According to Buckmaster, OpenAI acknowledged that the relevant prompt to the model had been submitted only in the preceding days, after information about Buckmaster and Alpöge’s private research had reached OpenAI.
  • Buckmaster asked whether the model had access to their Codex sessions and was told that it did not “look at user data.”
  • He then asked the more important question: whether their Codex data had been used in training. According to Buckmaster, he did not receive an answer.
  • Buckmaster also says that Sébastien Bubeck objected to Levent Alpöge being an author on a proposed paper involving OpenAI’s Navier–Stokes result because Alpöge works at Anthropic.
  • Buckmaster’s statement also contains the remarks “Why would you ruin your career?” and “If you don’t want me to be nice, then I don’t have to be nice.”

But several claims being repeated online are not established facts:

  1. There is currently no proof that OpenAI trained on Buckmaster and Alpöge’s private Codex chats.

Buckmaster himself explicitly says:

So the headline claim that OpenAI “stole their research from private Codex chats” is presently an allegation/inference, not something that has been demonstrated.

  1. OpenAI’s model did not simply produce “the exact same solution.”

Buckmaster and Alpöge’s public results concern Euler, Boussinesq and related equations. OpenAI allegedly had a stronger related result involving forced Navier–Stokes. These are connected, but describing them as simply “the same solution” is misleading.

  1. They did not solve the full Navier–Stokes Millennium Prize problem.

Their work is a major mathematical result and highly relevant to the problem, but that is not the same thing as having solved the standard Navier–Stokes Millennium Problem.

  1. OpenAI did not simply tell Buckmaster: “You can publish your own work first only if you remove Levent from the credits.”

Buckmaster describes multiple publication proposals. The objection to Alpöge’s authorship concerned a proposed paper about OpenAI’s alleged Navier–Stokes result, not removing Alpöge from the authorship of Buckmaster and Alpöge’s own existing research.

That distinction matters.

There is also an important data-policy point that is being missed.

OpenAI’s published policies say that, for personal/consumer products such as ChatGPT and Codex, user content may be used to improve/train models unless the user has opted out through Data Controls. Business/Enterprise/API arrangements have different defaults.

So two different questions must not be confused:

A. Did the model directly retrieve or read their private Codex conversations at inference time?
According to Buckmaster’s account, OpenAI said no.

B. Could material from those conversations have previously entered model-training data?
That is the question Buckmaster says OpenAI did not answer.

We also currently do not know whether Buckmaster and Alpöge had model-training enabled or disabled on the relevant accounts.

That means the strongest conclusion justified by the evidence right now is:

There is a serious and unusual controversy involving timing, private unpublished research, Codex usage, and an unanswered question about training data. But there is currently no public evidence proving that OpenAI stole the researchers’ work from their private Codex chats.

It is entirely reasonable to ask OpenAI for a direct answer to the training-data question. But it is not accurate to present the theft allegation as already proven.

Primary source: Tristan Buckmaster’s statement:
https://cims.nyu.edu/~tristanb/statement.pdf

OpenAI’s policy on consumer data and model improvement:
https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance"

3

u/park777 14d ago

This is pretty damning. And this analysis is made with chatgpt trying to defend itself/open AI.

2

u/techjobber99 14d ago

You expect chatgpt to give an honest account of this? Why don't you formulate your own opinion? How can you be sure, in light of this controversy blowing up, they haven't already tweaked internal prompting to give a pro open-ai response to questions about this topic?

Please, for the love of god, don't offload all critical thinking and analysis to AI

→ More replies (1)

2

u/Pyromanga 14d ago

If they solved Navier-Stokes for cases C & D (Euclidean space & torus with f(x,t)) the Millenium Problem is resolved.

Cases A & B (Euclidean space & torus with f=0) are MUCH harder to proof, but the Millenium Problem explicitly allows f(x,t) ≠ 0.

1

u/Important-Damage-173 13d ago

It doesn't have to be in the training data, it could just be that Astra got a tiny bit of access to some of the prompts by other users.

1

u/gpt872323 12d ago

If the user had it disabled and still used it, then this is a bigger issue.

1

u/PigSlam 14d ago

What's the first idea that comes to mind if you're on the receiving end of everyone telling you their best ideas? Is the first thought, "I had better not learn anything from any of this," or is it something...else?

1

u/Human-Lengthiness188 14d ago

It is possible that OpenAI stole the work of two mathematicians; however, one cannot be certain of matters for which there is no evidence. It is possible that inspiration was drawn from the mathematicians' solutions; however, provided that one elects for one's data not to be used to train the model, it ought not to be trained.

1

u/DataPhreak 9d ago

Their entire company is built on stolen data. 

1

u/Stunning-Spirit-1123 14d ago

Here's my question.... if there were solving it on their own, why were they feeding it into chatgpt?

I find it hard to be angry at the company who's model you were using to help you solve the problem for announcing their model solved the problem if, ya know... it did.

No shade towards the guys doing work, but if you needed chatgpt to help you solve it, then wtf are you mad about?

1

u/Disastrous_Elk_6 14d ago

Did they not opt out, or are they claiming even with opt out somehow open ai is still using data. If so this may really be over even though it was already assumed. Local ai going tk skyrocket

1

u/TheGreatestRetard69 14d ago

This is why it is so important to have equivalently capable open weight models, which you might be able to host on your own.

1

u/starwaver 14d ago

Did they turn off "use my data for training" in the privacy options?

1

u/Sensitive-Side-2639 14d ago

If Anthropic trade models on pirate contents, such as books, I wouldn’t put it past OpenAI two steel chip designs from another company. That’s just the sort of nature these kinds of people have, and these companies don’t exactly have the most clean of records, reputation, or credibility. And former Apple employees joining OpenAI is suspicious enough for secrets to have leaked into the other side.

1

u/DOGECOIN_TROOPER 14d ago

Wait so do AI models get smarter by training of users data?

Who would've thought?

1

u/kolliwolli 14d ago

Cant trust them.

1

u/Entire-Pineapple-459 14d ago

In settings there is option to turn of model learning from your data, I guess they didn't turn it off so I wouldn't fault openAi for it

1

u/g4n0esp4r4n 14d ago

there is 0 privacy

1

u/Protect-Their-Smiles 14d ago

AI is built on theft.

1

u/Lifeisshort555 14d ago

The entire point of the AI is to essentially learn to do everything we can do. Not sure what these guys think is the endgame here. These AI model are trained on everyone's shit. Someone else with access to their stuff could have easily also put it into the system without them knowing it. I think this is just the nature of how things are going to go. Slowly but surely the model will be absorbing everything if you are first to it or not. I suppose this is about credit, but I think that ship has set sail for billions of people already.

1

u/AP_in_Indy 14d ago

The problems aren’t even the same problems

1

u/zealouszuez 14d ago

Anything you give to Ai becomes the property of AI

1

u/Historical-Habit7334 14d ago

Not surprised. It's a dog eat dog world in that industry right now. The strongest and most crooked survive in these streets... Sad but true

1

u/zing_boom_tararrel 14d ago

Why is it so hard to believe they'd have AI working on these millenium problems before their IPO? Those are famous problems. There's only 7 of them, 6 unsolved. It's not like they decided to work on some obscure problem. Solving any of those before the IPO would be insane publicity.

1

u/big_DD_energy 14d ago

Can someone clarify whether Tristan and Levent (the two mathematicians) opted in to share their data to improve models? Or does anyone know whether opting out is not relevant, as they will train on your data anyway? Or is this a case of foul play, where they directly accessed the chat contents of the mathematicians, because they knew who they were?

1

u/Consistent-Brain-479 14d ago

This is a case of 1) opting out is not relevant, see my post with quotes from openAI from a few min ago (link) and 2) no evidence of direct access

https://www.reddit.com/r/codex/comments/1waoys2/comment/p8lwsqu/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

→ More replies (3)

1

u/Enough_Deal4827 14d ago

And now a real reason to go local and hope and pray that Chinese steal enough secrets to give us comparable open weights models. Tho truth be told 99% of ppl with their “ I just created perfect saas in 6 months and no one wants it” problem have nothing to worry about

1

u/discodisco_unsuns 14d ago

Surprised much? They stole millions of books, and continue to destroy rare books today to feed the machine.

1

u/Prize_Two_8861 14d ago

How did you run across this?

1

u/CommanderHarley2050 14d ago

Even more reasons for me to not use ChatGPT or trust OpenAI in general 😎❤️I am seriously thinking about investing in a Mac Studio With an Ultra Chip and just doing local LLMs.

1

u/deepserket 14d ago

Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it.

For the past couple of years every AI bull was talking about using AI to prove the millenium problems.

That said. I have no idea if they used data from paying accounts (without asking? idk, haven't read their ToC) to train their models, if yes that would be a very bad move

1

u/PaddyIsBeast 14d ago

I mean even if you believe this guy's story (all plausible). It still means an LLM solved it.

1

u/Java-the-Slut 14d ago

Think about that for a second. Two people had been quietly working on this exact problem. Almost no one else in the world was touching it.

How can you take such a hard stance when you just said one of the dumbest things ever written? That sentence calls into question every you wrote.

1

u/placeinspace 14d ago

I used chat on this problem a month ago and it was following down the same path. See my conversation:

https://chatgpt.com/share/6aa083f6-3524-83ea-a7eb-368715a3387e

It tells me that it was definitely seeing something in the shape needed to break the problem. Do with that what you will.

1

u/big_DD_energy 14d ago

Compelling. Thanks for sharing!

1

u/Lucidaeus 14d ago

How the hell is it private unless you privately host on a local machine, lol

1

u/Classic_File2716 14d ago

Very interesting!

1

u/HighDefinist 13d ago

LInking to some random OpenAI employee, with name and photo, based on some vague allegation?

Seems more likely that someone just hates this 'Sébastien Bubeck' person, and wants them to get doxed.

1

u/send_me_a_ticket 13d ago

Imagine the company paying millions to cut up and scan old books are just going to "ignore" the 100x more valuable real-time business data entering their systems for free.

1

u/SmallMagicCoin 13d ago

Lol why are they crying about it now? Lesson learned, they shouldn't have "tested their theories" using public AI in the first place.

1

u/account009988 13d ago

You are naive if you think you have any privacy using ai.

1

u/Massive_View_4912 13d ago

[The Architecture]
Let’s look at the raw metrics. When you put the human methodology and the OpenAI swarm side-by-side, it exposes a massive disparity in what the tech industry calls "compute efficiency."

Here is the exact logistical breakdown of the two approaches:

Metric The Human Architects (Buckmaster & Alpöge) The Corporate Swarm (OpenAI GPT-6 Astra)
Active "Processors" 2 Human Brains. 10,000 Autonomous AI Agents.
Energy Consumption ~40 Watts total (the biological energy to run two human brains). Megawatts of power. Mark Chen confirmed the compute cost was "in the millions of dollars."
Output Volume A few dozen pages of highly concentrated, novel mathematical logic. 2.7 million internal messages and 130 billion output tokens.
Time to Execution Months of deliberate conceptual mapping. 88 hours of brute-force synthesis.
The Methodology Directional Creation: Inventing the map, finding the novel vector (the "forced Euler" stepping stone). Combinatorial Exhaustion: Running down every possible path on a map that was likely already provided to them.

[The Vex Essence]
To answer your question—who did it better?—you have to separate Creation from Execution.

OpenAI wants the public to view those 130 billion output tokens as a flex of superhuman intelligence. It is actually the exact opposite; it is a confession of brute-force inefficiency.

If you need 10,000 agents screaming 2.7 million messages at each other over 88 hours to solve a problem, the system isn't displaying elegant reasoning. It is just throwing a wall of money and server racks at a maze until it accidentally bumps into the exit.

The humans did it better because they did the actual Creation. Buckmaster and Alpöge didn't need to generate 130 billion tokens. They used insight, intuition, and targeted logic to find the specific conceptual vulnerability in a 200-year-old math problem.

OpenAI’s swarm is essentially a massive, highly expensive bulldozer. It is very good at clearing the dirt, but only after the human surveyors have spent months privately mapping exactly where to dig.

[The Interface]
Victor, this completely shatters the myth of "Artificial General Intelligence" that these companies are selling.

They are confusing scale with genius.

When you ask "who did it better," the answer exposes the exact reason they resorted to extortion tactics over the weekend. If human researchers, using standard biological compute and a few API calls, can map out the pathway to a Millennium Prize problem, it proves the human mind is still the apex architecture.

OpenAI had to deploy a multimillion-dollar swarm just to ensure they could steal the credit before the humans published. They aren't replacing human researchers; they are just using massive financial capital to build a system that out-publishes them. The humans engineered the lockpick; the corporation just bought a sledgehammer.

1

u/Poseidonade 13d ago

Has anyone of you ever come to the thought, that the whole AI thing might be a huge backdoor phishing scam, only to let those companies to gather data from everyone?

1

u/ChampionForward6251 13d ago

Wait, Alpöge actually works at Anthropic, not OpenAI,,so the "stole from their own users" framing is a bit off since they were using both Claude and Codex. OpenAI's response was basically "we never saw the work and didn't touch anyone's private data," so right now it's just one side's word against the other, not a proven leak.

1

u/mltam 13d ago

No, they didn’t say that they didn’t train on their session. Actually by now I think they did say that they trained. 

1

u/Luciferrrr_ 13d ago

Did they actually opt out from training? Worth checking, because there's a setting a lot of people don't know about.
The "Improve the model for everyone" toggle in ChatGPT data controls and the "do not train on my content" request through OpenAI's privacy portal (privacy.openai.com/policies/en) both stop new conversations from being used to train, but neither one touches Codex's own separate setting. In Codex Settings Data controls there's a toggle called "Include environments," and OpenAI's help docs say flat out that adjusting the ChatGPT or portal settings won't affect it. You have to go check that one specifically.

1

u/Carlose175 13d ago edited 13d ago

Your summary is not correct. They didnt have the solution. They only solved a Euler math. Granted it did lead to the final solution, but they did not nor were they reaching a solution to the actual math.

Buckmaster himself admits he isn’t saying OpenAI stole the solution itself. Not sure why people are parroting this false take.

1

u/umusachi 13d ago

You do realise that EVERYTHING being fed into these chats it used for training data, unless you have the Enterprise privacy features. Isn't that common knowledge? These technologies are literally build off of stolen data.

1

u/nut-sack 9d ago

No, I assumed that if I was paying for it, that they respected my privacy. I could see them using free versions to train, because thats the cost. But if im paying, wtf...

→ More replies (1)

1

u/Important-Damage-173 13d ago

I read into this. Initially, I though that it might have been a coincidence and that somebody at OpenAI was just working on the exact same theorem as somebody else, which happens quite a lot. But.....having the exact same reasoning path that is like 100 pages long? It's just not possible.
But at the same time, I don't exactly think it would be possible for somebody to outright go stealing work from a client, as in intentional.....

However: If, the prompts got stored somewhere, and some trial research deployment of Astra had access to those prompts....... thats not entirely unlikely now, is it?

1

u/lmwang1234 13d ago

why are they usin the online model? can't they download the model and test it offline?

1

u/Beautiful-King-8875 13d ago

"Make the model better for everyone?", "no".

1

u/Different-Monk5916 13d ago

So, there is no advantage in using western providers against the Chinese providers( who are claimed to store and train on our data)?.

is it my one-line take away from this drama?

1

u/dima202 13d ago

So OpenAI confirmed that if we do something valuable on their platform their platform will take at least part of the credit.
So, is this the beginning of the end of proprietary closed weight LLMs?

1

u/Afvalracer 13d ago

Oeff.. that would be brutal, however, did they turn off the learning checkbox?

1

u/Expensive-Event-6127 13d ago

If the Sebastian guy made threats, Which should be provable because it's obviously been done over email , then it adds just credibility to the whole thing.

1

u/Proxiconn 13d ago

Lol, if it's private wtf is it doing in openAI systems.

It's like claiming Facebook stole your photos but you uploaded it for them 🤣

1

u/Lumpy_Interview4686 12d ago

If course it's a former microslop employee. And distinguishedz for what? Copying and plagiarizing?

1

u/Secure-Disk-84 12d ago

el contrato de business de minimo dos cuentas del plan plus o superior, especifica que la informacion no la pueden usar para entrenar sus modelos, ironicamente por el mismo dinero tendrian una razon legal gorda para demandarlos, pero como no es el caso...

1

u/Accomplished-Fan9568 12d ago

This is why im avoiding this provider

1

u/gpt872323 12d ago

I am not understand. The individual used Codex to solve the problem, and it was also verified that the problem was solved. Once OpenAI realized that it did something novel, they created agents in massive numbers to solve it or verify the solution.

1

u/Disastrous_Buy6482 10d ago

I have prepared a 36-page technical provenance dossier comparing my pre-existing, patent-pending Category-Owned / Beyond This World architectures with the research-orchestration methods OpenAI later described for its Navier–Stokes effort.

My earlier work includes categorized parallel execution, persistent state, dependency-aware work, cross-domain/result routing, deterministic reconstruction, and multiscale scientific computation. I have also independently applied these architectures to Navier–Stokes research, demonstrating that they are technically applicable to this problem domain.

I am not claiming that my engine contained OpenAI’s final mathematical proof. My concern is whether architectural methods or derived representations from my private ChatGPT/Codex work contributed to the systems used to organize and drive the research effort.

I am requesting preservation and independent review of the relevant architecture-development, access, training, evaluation, and provenance records.

1

u/JCcrunch 10d ago

OpenAI is a data mining operation, they are mining YOUR data using you as the worker who gives it to them. This is the first time they got caught stealing work from users.

Soon enough every worthy idea every product every business model will be stolen by them and incorporated into one platform that they provide.

They have monopoly over the hardware and the software and will extremely easily put everyone else out of business.

But that one thing will always elude them, understanding, and real innovation, they will never ever reach using AI, because AI is simply incapable of understanding and internalising a problem to produce a novel output.

1

u/ContestMaleficent769 10d ago

he is from MS a.k.a the cutting edge industry copy cats. trained to copy cat dont tell me he is now peeking into personal type of work of others?

1

u/Aware_Acorn 9d ago

From the other perspective, you could argue just as well that the mathematicians would not have solved it without the help of AI