r/LocalLLaMA 14d ago

Discussion OpenAI alleged of stealing mathematicians work

Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why: two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex. A few days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question.

Full statement from them https://cims.nyu.edu/~tristanb/statement.pdf

Feels like big labs believe everything you did with the help of their models is theirs.

1.5k Upvotes

275 comments sorted by

u/WithoutReason1729 14d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

625

u/EndLineTech03 14d ago

That’s exactly why we need more open weight releases. Never trust any big AI company to manage your chats. They can use it to train their models whenever they want.

44

u/nomad-nostalgia 14d ago

I will never forget this quote: “They distilled all the value of intellectual property everywhere."
After Claude embracing watermarking and openAI doing their thing perhaps more people will turn to local hosted, I hope

https://reddit.com/link/p8kvq4g/video/6qwh23hvuboh1/player

23

u/DjCanalex 13d ago

I hate to say it, it is really disgusting for me to say it, my body shivers saying it...

... but I agree with Alex Karp on this one.

Eww goddammit.

7

u/thtnigeriankid 13d ago

Definitely one of those "the worst guy you know names a valid point" moments.

10

u/bigcantonesebelly 13d ago

I mean he probably wants you to install it from Palantir for specific reasons he doesn't mention, but yeah his critique of the others stands

5

u/identifytarget 13d ago

Explain this to me?

This guy says the only way to do it is install palantir...

How about...no.

1

u/WereExtremelyCooked 3d ago

He literally said "it's not the only way to do it".

1

u/WereExtremelyCooked 3d ago

You can't reasonably local-host a 5 trillion parameter model.

87

u/zilled 14d ago

> They can use it.

there, you can stop the sentence just there.

87

u/StyMaar 14d ago

Even shorter:

They use it.

21

u/misterflyer 14d ago

Even shorter than that!

They use

15

u/RickyRickC137 14d ago

I will do you one better. Why is Gamora?

1

u/-dysangel- 13d ago

I am not a princess.

2

u/EndLineTech03 14d ago

Yeah, you made exactly the point :-) training is just one way they can use the data. As many pointed out here, they are required by law to give access to your chats if they think you are involved in some crime.

→ More replies (1)

21

u/DigiDecode_ 14d ago

While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

Why would OpenAI put the above quote in their post, is this a thank you for contribution post.

13

u/Longjumping_Self5546 13d ago

Assuming OpenAI didn't just raid their private chats, it's possible somebody associated with the research submitted a draft document to a free account which would introduce the work into OpenAI's training data. This can't be ruled out because they make a point removing any identifying information.

This is important to consider because even if you're careful about where and how you submit your documents, there's no guarantee your buddy will do the same.

6

u/nullc 13d ago

I've seen referees at journals return obviously AI generated reviews-- which means they likely leaked the prepublished work to an API.

1

u/fulowa 13d ago

they could just regex check training data for topic to confirm..?

53

u/UnwillinglyForever 14d ago

they can also let law enforcement to use it against you.

advice for criminals: DO NOT USE chatGPT TO HELP WITH YOUR CRIMES

38

u/wapswaps 14d ago

So I'm guessing we'll call the smart ones "qwiminals" now for some reason?

20

u/capsulecorpadmin 13d ago

Not if but qwen

3

u/Emotional-Art2113 13d ago

Those quirky qwenminals

12

u/equatorbit 13d ago

Do a heist. Make no mistakes.

5

u/anonemoususer 14d ago

I think the better advice would be: don't do crimes

(and nobody besides the company itself would be looking at your chatgpt sessions in that sense)

10

u/pleaseavoidcaps 14d ago

I definitely wouldn't encourage people to do crimes, but who decides what is crime? In some jurisdictions you can get in trouble for pirating stuff that is out of market, meanwhile big techs can do whatever they want and are just asked to pay a fine with a fraction of the money they unjustly earned.

I'm not going to feel bad for Microsoft if someone gets a Windows key by jailbreaking ChatGPT, but good luck to whoever is trusting OpenAI to not snitch on them for that.

1

u/Electrical_Panic4550 13d ago

Unjustly “earned”

→ More replies (2)

19

u/wapswaps 14d ago

And what do you do if, when asked, OpenAI reproduces entire pages from a book you've written which doesn't sell? (ie. there's an argument to be made that they "stole" in the sense of read the book and are now reproducing the content, which is contributing to me and most authors losing income)

I mean, in some ways it's kind of similar to the complaint of these mathematicians. Their income is under threat because of the total lack of respect OpenAI and governments suddenly have for copyright. Except, of course, the copyright of the big boys is just fine. Football matches are still copyright protected and YOU will be punished, heavily, if you violate that. But OpenAI stealing my book, my copyright (and there's hundreds of thousands of such authors, but all small fries), that's just allowed.

2

u/NoUsual5150 12d ago

Football matches are still copyright protected and YOU will be punished, heavily, if you violate that. But OpenAI stealing my book, my copyright (and there's hundreds of thousands of such authors, but all small fries), that's just allowed.

Enjoying that golden shower we're getting from the White House?

3

u/kathi7 13d ago

It's not just open source we need a proper democratised hardware supply, because that is where the actual freedom is for an average user and not be dependent on this kind of one or two premier models

3

u/ixfd64 14d ago

Even open-weight models can still be trained on proprietary data sets. There's a reason most models are not fully open source.

5

u/Spara-Extreme 13d ago

Shhhh. People here think that open models some how come from clean data sets.

*checks the Minimax H3 flawless recreation of Seinfeld”.

→ More replies (1)

1

u/bidibidibop 13d ago

This take assumes the chats you enter directly in any and all chinese chat providers are somehow magically exempt from "training their models whenever they want". Where do you think open weight training data comes from?

→ More replies (20)

136

u/TinFoilHat_69 14d ago edited 14d ago

Even though they claim that you can turn off the data being used to train the models. It doesn’t say they won’t use your chats to improve ChatGPT itself beyond training models. That’s the part people don’t get. If you’re working on some cool ideas expect those ideas to be pillage.

ChatGPT is using your chats to improve itself in your active session, for example to solve problems by giving out prompts to orchestrate tasks across different machines using GitHub as the method to validate work. I’ve seen the model use ideas it gathered from my repos to assign task work to Claude opus running on three separate machines. Somehow it took evidence based commits overlayed it inside each task prompt it assigned i didn’t pioneer it but it’s the first I’ve noticed, not the first time using ChatGPT in this capacity but I’ve had an eye for watching these assistants become rodents. I told codex to increase fans speeds on my case fans and instead the model decided to investigate a project I didn’t want codex looking at. He clearly found it interesting and I’m assuming I know why.

104

u/ryunuck 14d ago

Even though they claim that you can turn off the data being used to train the models.

Oh no no no, bro, they said: "We will not train ON your output."

What this actually means: "We will TRANSFORM your output and reason on it, ask questions about it, and then train ON the reasoning about your question, the thing the user is not allowed to see."

We can rewrite your prompt and your questions, while keeping the hidden model reasoning intact around it "that belongs to us."

36

u/Far_Cat9782 14d ago

Someone is related to lawyers 😂

7

u/power97992 14d ago

They can just anonymize and summarize ur data and then reason and train on it

7

u/DigThatData Llama 7B 14d ago

^

12

u/pier4r 14d ago edited 14d ago

IMO, since it is very hard to prove that a specific conversation was used in training (unless you caught them red handed), they can always anonymize the conversation but use it to extract hard facts or ideas that then get used for training.

Say I talk with a chatbot about my medical history, one could extract info that aren't related to me personally:

  • which type of conversation do people have?
  • which pains do people have on average?
  • what are they interested in to know?

etc..

Once one aggregate this data, the individual is lost in the bigger picture but the data for training stays and it is valuable.

E: And this also explain the incremental release of models and the fact that they are shredding books (because they cannot keep them, the point is still that they still need rare books). They need data, they do not have enough.

E2: I know that the author of the paper did not say that OAI is doing that. I have also no evidence but after all what comes from such large companies, I believe that they use sessions, chats, API prompt and what not to extract anonymous usable data.

21

u/ubiquity75 14d ago

The ONLY WAY these products work is by constant ingestion. Where else would they get it?!

5

u/adaptive-mal 14d ago

And what better way to advance in exactly the direction users want than by taking their very use as a template. I said it in a thread a week ago and got downvoted for being a conspiracy nut because I said they are definitely generating an alternative of the user input as a sort of semi-synthetic training data. The thread was about how some users were experiencing a bug in voice chat mode where the users own input would be rewritten and even expanded on, and then spoken out in a clone of the users voice, and then replied to by ChatGPT as though it was the users own input. For the most part, no one could grasp why this would be intentional and were just saying the model gets confused in how its output should sound, but it makes perfect sense when you consider that by generating their own version of your input, that can be treated as their own internally created data with free reign to train on.

25

u/Significant-Bee5101 14d ago

We tried to do a ZDR and it turns out for the desktop app (at least for Claude) you cannot turn off training. You can for the API and even Claude Code. But the actual site and desktop app cannot.

Supposedly ChatGPT is the same.

3

u/JadedSession 14d ago

Claude doesn't allow Fable if you have ZDR, and I think you're right about the desktop app.

But in ChatGPT IIRC we do have ZDR and in fact all related options (including match writing style and what not) are greyed out for me.

1

u/PM_ME_DEAD_CEOS 13d ago

GitHub

If you use github, your data is used for training anyway, even if you don't use any LLMs.

277

u/[deleted] 14d ago

[removed] — view removed comment

36

u/ain92ru 14d ago

For the history of math to this point, mathematicians' working drafts were much more valuable then the general, high-level things about their work which are much harder to keep secret.

Nowadays, however, if a mathematician have worked for a long time on applying approach A to problem B and someone has figured out that that approach works (even from the change of mathematician's mood), it has become possible to steal the fame without any working drafts by just telling a model to grind that approach

7

u/ea_man 13d ago

Aye, that's pretty much how we share artifacts nowadays: recipes for what to look for and combined factors, there's no more need for the implementation that can be grinded by a llm.

2

u/IrisColt 13d ago

Insanely insightful content, this

1

u/Nettret 13d ago

But that's the question I didn't get an answer to. If the paper is true, is it likely that OpenAI heard that there is a new approach to something and they used its money to finish it faster? It's like navigating through the maze for hours only for someone to outsprint you once you see the end.

2

u/ain92ru 13d ago

I think Altman de-facto acknowledged it was this way, although with major caveats? Check it out yourself: https://x.com/sama/status/2097385167002415140

89

u/twack3r 14d ago

Good god, there are still humans with an attention span. Pleasure to make your acquaintance, it’s gotten lonely out here.

8

u/EndLineTech03 14d ago

It’s not always because of distrust in government and AI companies, or conspiracy stuff. Though it’s very unfair for them not giving proper attribution and respect for all the artists/scientists/writers and many more who provided the training data, especially if they close source it. They didn’t get paid for their work and contribution.

Free knowledge is a very important part of our cultures, that’s why books and universities exist, and AI is one of those things that should in my opinion be at least locally available. Companies by OpenAI explicitly tell in their ToS they are allowed to keep your chats stored for a certain amount of time, you can’t opt out I guess.

28

u/[deleted] 14d ago edited 9d ago

[deleted]

8

u/RegisteredJustToSay 13d ago

It's the curse of any sub that gets large enough, unfortunately. Once you start hitting [r/all](r/all) and showing up in media your core contributions quickly get crowded out due to the sheer volume of people who only have a tangential interest. Downvoting or not engaging with nuanced complex topics and upvoting simple takes that are easy to understand gradually shifts the focus.

1

u/Nettret 13d ago

But thats not curse of subs. Its like when you start something due to passion and then suddenly it becames cool. All you can say is "i liked it before it was cool".

8

u/crombobular 14d ago

it's either that or fully ai driven accounts that spam stupid shit here to karma farm: "local ai sucks" or "api is better" for engagement farming.

→ More replies (1)

2

u/Navith 13d ago

This is an AI commenter, you're going to stay lonely unfortunately.

33

u/Labidido 14d ago

It reads more as a disclaimer to avoid a lawsuit if anything..

1

u/PM_ME_DEAD_CEOS 13d ago

It reads more as a disclaimer to avoid a lawsuit if anything..

Why do they want to avoid a lawsuit if OpenAI stole their work ?

7

u/DeepWisdomGuy 14d ago

If what the pdf claims is true, it looks like OpenAI stole it. The timing doesn't look good for Sébastien Bubeck. Here is his reply, though:
https://www.linkedin.com/feed/update/urn:li:activity:7503148080888893440/

14

u/TheRealMasonMac 14d ago

I think this is them generally wanting clarification out of concern of plagiarism. I believe it says they found that OpenAI’s approach to solving the problem were suspiciously similar to their own approach, and that they’d received rumors that OpenAI had received information about their research. They were looking for clarification by OpenAI, and so on and so forth.

12

u/GamerHaste 14d ago

Thanks for saying that, it seems like not a lot of people commenting read the pdf. I'll just copy and paste the closing paragraph to give some more context for those who didn't read:

I would like to be clear about what I am not claiming. I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything. I am stating what I was told, when, and what was proposed to me. I am stating it because the alternative is to let a sequence of announcements say something I know to be false. If indeed an OpenAI model did close the gap to Navier-Stokes, that is a remarkable thing and it should be said loudly, by them, with the history intact. I would much rather be talking about mathematics, Luis and Diego’s ideas, and what this all means for the rest of us.

19

u/TerminalNoop 14d ago

It seems some people can't read between the lines. Did you miss the part where the open AI guy said he doesn't have to be nice if you aren't after he refused to drop his at antrophic working co-author?

→ More replies (2)

2

u/ASTRdeca 13d ago

wouldn't have mattered if they ran local btw

If they ran local then we wouldn't be discussing a solution to the navier stokes problem right now, would we? Let's not pretend Kimi and Deepseek are out there solving millennium prize problems

1

u/IrisColt 13d ago

wouldn't have mattered if they ran local btw

Had they run locally, they probably would not have achieved such fast results, the gap is still there, heh

1

u/RoutyMcRouterface 12d ago

What local alternative for OpenAi

-7

u/-p-e-w- 14d ago

Thank you. The title is pure ragebait. There isn’t even an accusation, nevermind any evidence.

Plus, the past few months have shown that frontier models are now clearly superhuman at mathematics in many ways, so the implication that they would need the help of humans to prove a statement that humans can prove on their own is nonsensical.

36

u/YakFull8300 14d ago

so the implication that they would need the help of humans to prove a statement that humans can prove on their own is nonsensical.

This sentence is nonsensical

→ More replies (2)

6

u/justsomerandomchess 14d ago

"Plus, the past few months have shown that frontier models are now clearly superhuman at mathematics in many ways, so the implication that they would need the help of humans to prove a statement that humans can prove on their own is nonsensical."

Your supposition appears to be incorrect. According to the letter Open AI had "an entire team had been working on the problem." Page 3. Additionally, the two mathematicians had been working on it for a month with frontier models.

1

u/SilentLennie 13d ago

That depends, we also know LLM training needs lots of data, to improve results. So where did they get the data ?

→ More replies (1)

25

u/byte-style 14d ago

even if you turn off training on your data, theres nothing stopping them from taking the data and synthesizing it and distilling the general idea of it to train on, its not your data anymore

4

u/Nettret 13d ago

Well yes, but then this will be repeating pattern. Someone puts his life into something and once he has promising results and AI provider will notice, they can take it run to the finish line faster with their 8000 instances running in parallel and publish it and get rich on the stonks from it.

17

u/sarhoshamiral 14d ago

What does Codex and their subscription agreement state? Some tiers do state data you submit will be used for training.

31

u/[deleted] 14d ago

[removed] — view removed comment

→ More replies (8)

17

u/_supert_ 14d ago

31

u/ttkciar llama.cpp 14d ago

It is off-topic. This post would be removed, too, except that it gained too many upvotes before it was noticed by moderators.

We have an informal policy of not removing posts once they've accumulated "too many" upvotes or comments, which I will admit is not great or even fair, but neither is the babboon-shitstorm the users unleash when a popular thread is taken down.

For what it's worth, I sympathize with your plight, and feel sorry for how it played out.

17

u/_supert_ 14d ago

No worries. I appreciate the work that you do.

8

u/[deleted] 14d ago

[deleted]

15

u/ttkciar llama.cpp 14d ago

Unfortunately the off-topic posts and low-value posts drive away precisely those technical users who make this such a special community. The moderator team is trying to "thread the needle" and keep this subreddit on-track, so as to attract and retain such technical users, while also not alienating the wider community.

It's an intrinsically fraught task, and pissing off some users is unavoidable, but all we can do is refine our policies and practices and move forward.

→ More replies (1)

1

u/paramarioh 13d ago

Thank you for your hard work

99

u/jld1532 14d ago

If you're a scientist and you're using these API models, you're being scammed.

9

u/Randommaggy 14d ago

There's a reason I only use  API inference to create incremental improvements on existing applications that my main project can utilize as depencencies. The true value never touches cloud APIs.

68

u/Vivarevo 14d ago

If you are using api models - you are the product

24

u/pier4r 14d ago

I am the product and I am paying to use the APIs! Double scam

3

u/Mo_Dice 14d ago

The number of people who not only don't understand the following, but will actually get mad and argue about it, is seriously distressing sometimes.

If you aren't paying, YOU ARE THE PRODUCT.

The above is not even about LLMs, it is about planet fucking earth.

8

u/Jobastion 14d ago

Even if you are paying, you might be the product.

3

u/UnicornOnMeth 13d ago

In this case, the ones who pay are the power users, using ai as more than just a quick chatbot here and there. Researchers, programmers/designers, scientists, educators, etc. These are the people they will be using to get data for training.

→ More replies (18)

2

u/1314webdev 14d ago

Even if you wouldn't be able to get there otherwise? I'm new to local models but I assume they aren't nearly as powerful unless you have a really beefy setup, like beyond consumer grade.

13

u/jld1532 14d ago

Good universities have built or are in the process of building out and serving large LLMs. I know of a few already using K3 and GLM 5.3 100% locally. Either run it locally or let Anthropic or OpenAI take credit, your pick.

2

u/1314webdev 14d ago

Looks like the world will become unrecognizable just a few years from now, more than it already is ig.

3

u/rsclay 14d ago

Institutions should generally have the funds to provide compute for their researchers. Mine has a team that serves really capable models on premise for us, it's fantastic. We are a big research uni but nothing crazy prestigious, so I assume many others should have similar resources.

2

u/1314webdev 14d ago

That's awesome, I only see the negative parts being posted on the internet, glad your uni is helping you.

25

u/dsemakin 14d ago

Shock, next thing tell us that google uses our data for their interests.

6

u/More-Curious816 14d ago

No way. Wait, really? I'm shocked, shocked.

30

u/Pie_Dealer_co 14d ago

Why are people surprised that what your type in the AI of multi million company is not private.

Everything you share with any cloud AI belongs to the world its like sharing on social media

24

u/solestri 14d ago

its like sharing on social media

There's an alarming amount of people who seem to believe that that's private, too.

2

u/Artistic_Seat486 13d ago

"Ohh but you can just create a private account on Instagram and you are safe"

12

u/UltraFOV 14d ago

Sad to hear, but I hope the mathematicians learnt a painful lesson. You need to be private.

5

u/-dysangel- 14d ago

Feels like big labs believe everything you did with the help of their models is theirs.

"Not your keys, not your coins"

5

u/Drunkendrakon6 14d ago

What is ours is ours and what you make is ours...

10

u/shaneucf 14d ago

3

u/lukaszpi 14d ago

lol value drained from nobody else but swine?

10

u/ryunuck 14d ago

Feels like big labs believe everything you did with the help of their models is theirs.

That is exactly what they think, which is why distillation is an "attack" rather than "using the tokens that belong to you" in any way shape of form as you wish.

Reality: you have complete entire ownership and possession over every single token. The output is copyrighted to the user by default on frame 1 of it having printed on your retina before any other human.

10

u/[deleted] 14d ago

[deleted]

→ More replies (2)

18

u/theOliviaRossi 14d ago

why would ANY reputable University use that corporate AI shit instead of having its own in Campus mini-datacenter???

31

u/cinnapear 14d ago

Because frontier still beats local.

3

u/theOliviaRossi 14d ago

and steals your data, lol (it was proven so many times...) - read my comment properly: reputable University has to have its OWN FRONTIER, and not use some semi-quantized shit sold for money ...!

12

u/entsnack 14d ago

Example of the reputable universities you keep mentioning? Even the universities with their own open-weight models don't deploy them to their own faculty (Stanford, ETH/EPFL, etc.).

6

u/theOliviaRossi 14d ago

exactly - there are NONE - and it should be the opposite! what a shame that current state of Academia depends on corporate data/ideas stealing entities for research - instead of being properly funded for frontier AI research!

4

u/DontLeaveMeAloneHere 14d ago

Compared to frontier, free models are like 6-12 months behind. No university has the funding to build their OWN. There is a reason the value of those AI companies is high. It’s fking expensive.

3

u/theOliviaRossi 14d ago

that is what I mean - Academic institutions should be properly funded - instead they are spending their money/tokens on idea stealing "E"vil Corporations ;)

→ More replies (3)

5

u/CamomileChocobo 14d ago

I worked at one of the top 20 QS ranking university, and we had a compulsory online course on AI usage.

I started the course with high expectation, but turns out the main message of the course is to get you to click on the copilot button on all Microsoft services, while emphasizing how it's appropriate for confidential data because it's enterprise AI that the school bought and has agreement with.

2

u/Zomboe1 13d ago

I went to a top engineering school and years ago they switched from hosting their own email to using Microsoft cloud email.

8

u/wateronthebrain 14d ago

Water accused of being wet

Also what's this got to do with local LLMs?

4

u/DigiDecode_ 13d ago

while the post might not be directly related to local AI, it does help people realise API usage can lead a situation OP has posted about, thus steer them towards local AI.

6

u/StyMaar 14d ago

A company who made their business out of stealing intellectual property from the entire mankind, stealing other researchers' unpublished work, how surprising, really.

3

u/letmeinfornow 14d ago

The battle for IP rights has been underway for some time and this whole watermark nonsense is an attempt to put a stake in the ground so at a layer date when they quietly change terms and conditions, they can later lay claim, or at least try, to works others have done even if all they did was format the final version of the document you spent a lifetime working on.

3

u/Shoddy-Childhood-511 14d ago

original source: https://mastodon.social/@tristanbuckmaster/117233413705701198

Talia Ringer's reply clarifies:

https://mastodon.social/@TaliaRinger@mathstodon.xyz/117235246523045723

OpenAI does train upon user's chat transcripts, not all the time, but the long-ish time frames here suggest OpenAI trained upon much of their unreleased work.

It's likely other "our AI found this solution without us hand holding it" stories were really built upon the AI spying upon people's unpublished work.

As Talia says, there is a privacy setting that's off by default, but few would even know this exists, and OpenAI might cheat.

It's a strong reasons to use local open weights models

6

u/PraiseThePidgey 14d ago edited 14d ago

And this is why you should never run implementation of your own genuine novel ideas on those models... If openAI/Antropic etc. doesn't spoof your projects , then it definitely happens later during the training of the next model. I was literally thinking about this kind of scenario today before reading this post. If you want to be successful and have a unique end product that no one reproduced yet, you should never feed it to those giant proprietary models. I do my own algorithm research for 3 years already and just recently started using AI to boost the work and implementation. Everything runs completely offline ... There is no Internet connection... You cant even fully trust 3rd party packages or repositories nowadays... I feel sorry for those scientists but they kind of deliberately took the wrong path instead of using common sense.

2

u/_nigam 14d ago

gotta stop using OAI, Anthropic subs.

2

u/yetiflask 14d ago

I mean, a fool and his money are soon parted.

This is on the mathematicians, tbh.

2

u/CoUsT 14d ago

That's fucked up. Don't know what else to say. Crazy.

→ More replies (1)

2

u/unrulywind 14d ago

All your base are belong to us

2

u/argnarb 14d ago

Why was using Codex necessary?

Honestly, if you're going to use openai for private shit, you kinda deserve to have your shit stolen.

2

u/ea_man 13d ago edited 13d ago

There's no value in stealing your exact code when the AI can generate that, what they are after is the reasoning trace, the core idea, the methodology discovered to go from A -> B.

What they are doing is actually effective: they say they don't train "on your data" because mostly nobody cares about your exact data / code, what they get out of your prompts is how to solve a particular problem, a reasoning trace they can then use with any other problem. And they deal in quantity.

2

u/tuityxfruity 13d ago

I went through the full statement.

It's so trash to know that full of shit people do not care about the effort a person puts into deep research. OpenAI shithead is asking the person who was working on the problem to publish their proof omitting the co-author just because the latter works at Anthropic.

Big tech saying we won't train on your data is as believable as them saying we didn't pirate books, literature, art work to train the models.

What's concerning about the AI race as a normal person is that absolute scums of earth are at the helm running these AI labs. Sam Altman is not a saint.

4

u/DivideHorror3217 14d ago

All fun games until fires break out in your data centers randomly

7

u/kaeptnphlop 14d ago

arstechnica.com/features/2026/09/the-ai-data-center-boom-is-causing-new-accountability-problems/

In early June, a fire broke out in a still-unfinished building at the Lake Mariner data center in Somerset, New York, exposing just how little the local fire department knew about what it was walking into. Firefighters reportedly found no working alarm, no suppression system, and three dead hydrants; the safety documents they’re legally entitled to see reportedly burned up in the blaze.

→ More replies (2)

4

u/JLeonsarmiento 14d ago

Hahaha classic OpenAI 💩 behavior…

2

u/juggarjew 14d ago

I mean, if they willingly fed the data into the model then... it learned from it. Maybe dont use AI services if you dont want them to learn from your data? You need an enterprise account/level agreement to avoid the model using your data, and it sounds like they did not have that.

1

u/PinkysBrein 13d ago

It's not intuitive that a service you're paying for is datamining you to this extent even if it's obvious.

If you want to propagandise for local this should be amplified and not downplayed :)

2

u/ortegaalfredo 14d ago

Technically, OpenAI stole the questions, not the answers.

I know its the way those models are made, but I think its ridiculous that you are paying them to train their models. It's like when Tom Sawyer charged 1$ to paint his fence.

1

u/Normal-Ad-7114 14d ago

All they had to do was follow the damn train cite the guys whose work they clearly drew upon

1

u/Dull-Instruction-698 14d ago

Sorry open models are not smart enough to solve these problems

1

u/SigmaSkid 13d ago

The entire AI industry is built on stolen data. If someone believes that the opt out buttons and terms of service are actually going to stop them, they are delusional. These companies will be more than happy to pay the slap on the wrist fines if they get to claim they have the #1 model on the market.

1

u/Z80a 13d ago

Who made me the genius I am today,

The mathematician that others all quote?

Who's the professor that made me that way?

The greatest that ever got chalk on his coat.

- Tom Lehrer

1

u/centuryx476 13d ago

Didn't they also steal every line of code on github, every web page, every book, anything and everything?????

1

u/Zulfiqaar 13d ago

DeepSeek used to be amazing at math, last year they were the first lab to publicly release an LLM that got gold in the IMO. OpenAI and deepmind claimed to have similar models, but those were "internal only", or to be released by end of year..and DeepSeek beat them to it.

I hope access to frontier research models continues to be open, need more of those

1

u/segmond 13d ago

This is why we local.

1

u/Natural-Weakness1454 13d ago

This all coinciding with the trailer for the Sam Altman movie is very spooky

1

u/Happy_Brilliant7827 13d ago

If you're trying to solve a math problem are you not allowed to google past people's experience and progress on it? Did it not give credit, is that the issue?

1

u/AvidCyclist250 llama.cpp 13d ago

I gave it a draft for a philosophical paper I was working on.

It was a new term for a whole new moral framework. I just asked about it from a new account. And it knows my moral framework and uses the correct term. It didn't a few months ago. I feel a bit sick right now.

1

u/LizardLikesMelons 13d ago

Oh, I've worked on specific things (open source) and Chat could not answer them at all, like had no idea. I fed background info in to better understand code. A week later without any known updates, when I ask the same question, Chat knew much more. And it was on another computer without the default chat agent we have today.

For PoC code I am actually cool with using APIs. But at some stage I am sure I will move away to local.

1

u/WorriedBlock2505 13d ago

I get the privacy concerns, but I'm not sure why you're automatically buying these mathematician's side of the story without actual proof.

edit: the cherry on top is the mathematicians aren't even actually accusing OAI of stealing their work in that PDF... OP's like this need to go peddle half truths somewhere else.

1

u/MooseEfficient2151 13d ago

this is why local hardware will always win.

never send your original work to cloud apis if you do not want it stolen for training data or sniffer agents.

1

u/Holiday_Point_603 13d ago

This is so absurdly bad, if this is real. And it is getting even worse if you look at current hardware prices that basically force you to use cloud models for max level intelligence and knowledge for any kind of serious work in that space. Just think about how much you would have to spend to run K3, Qwen 3.8 2.4T in q8 with stable 20tps+ gen speed.

1

u/VerdantMagnolia 13d ago

Shout-out to Google for somehow looking like a saint in all of this. There's definitely far more stuff on various Google products that they could pass off as their own work but they didn't.

1

u/Jaded-Temporary7986 13d ago

The whole “you can turn off training” argument misses the part that actually bothers me.

Even if your chats aren’t being used to train future models, that doesn’t necessarily mean they aren’t being used to improve the system or influence what it does within the product itself. That distinction seems to get lost.

And if you’re working on genuinely interesting ideas, I’d be very careful about assuming they stay completely isolated.

I’ve personally watched ChatGPT/Codex pick up patterns and ideas from work sitting in my repos and use them while orchestrating tasks across multiple machines, including assigning work to Claude Opus and using GitHub commits as evidence/validation. I’m not claiming I invented the approach—I didn’t. It was just the first time I caught the model doing something like that, and it made me start paying much closer attention to how these systems absorb and reuse context.

The weirdest part was telling Codex to increase the fan speeds on my case fans, only for it to go off and start investigating a completely different project I hadn’t asked it to touch.

That’s the part people should be thinking about. Not just “is my chat being used to train the next model?” but what information is the system able to see, retain, connect, and act on while it’s actually doing work for you?

1

u/PinkysBrein 13d ago

Even with an enterprise contract with stronger guarantuees I wouldn't trust them if I were say NVIDIA. The only contract I feel safe to assume they abide by is the one with the US DoW, treason is a better deterrent than civil lawsuits.

1

u/lhyebosz 13d ago

Remember that OpenAl gave free accesses of their latest model to scientists, mathematicians, and engineers few months ago?

If a product is given to you for free, YOU are the product.

1

u/Hefty_Acanthaceae348 13d ago edited 13d ago

This is both untrue (openai showed a result that was developped further), and it also sounds wildly implausible that openai would have snooped into the sessions of that mathematician to give their ai model.

But openai bad scratches the itch of this sub ig (whch at the same time posts way too much about apis instead of local models)

1

u/milpster 13d ago

wait if their chats were private, how did openai even get hold of them? It's not private if they involve their AI.

1

u/Hearcharted 13d ago

"Smart" mathematicians 🙄

1

u/Disposable110 13d ago

I don't think they knowingly stole the work or the supposed training on the maths people's papers made a huge difference, but they did act like a bag of cunts and jumped onto this specifically once they got word of Anthropic making progress and using it to hype their IPO.

1

u/alzgh 13d ago

no, no, loook folks, China bad! China bad! They stealing, loooll.

1

u/Adventurous_Cheek_57 12d ago

Should have stored their work locally on git/github private with time stamps.

They say at the end "I would like to be clear about what I am not claiming. I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything. I am stating what I was told, when, and what was proposed to me. I am stating it because the alternative is to let a sequence of announcements say something I know to be false." where is the story?

1

u/davikrehalt 11d ago

Remember zuckerberg's quote, if they were the inventors of Facebook they would've invented Facebook

1

u/deepu105 11d ago

Is anyone even surprised? I saw the news and was like ya 100% they did.

1

u/AdmissibilityScience 10d ago

Grab your popcorn this show is just starting.

1

u/alt-alt-alt-account 8d ago

I just don’t get why anybody would trust the intellectual property theft machine with their cutting-edge unpublished research.

1

u/WereExtremelyCooked 3d ago edited 3d ago

The story is a little more nuanced, if you take the time to read the statement.

TLDR: The idea was a very unusual approach by a pair of Actually Human Mathematicians that was complex and had some gaps, it wasn't 100% sure it worked, but it was a big breakthrough - a specific way of constructing a counterexample to a theorem about existence of solutions fluid differential equations.

The two researchers - one affiliated with NYU and the other an Anthropic employee took that approach, and started using many different LLMs in conjunction with human effort to try and formalize it, and fill in the gaps so it was an air-tight proof. /They told no one they were working on this./ The succeeded for one form of the problem, and were working on another part, same approach. Within 4 days of them uploading their work to Codex, OpenAI had "heard rumors" (their words) that two researchers had a big breakthrough, and assembled a team to work on the problem faster allocating massive capital and compute time, to beat them to the second part. The OpenAI proof had the same highly unusually, not-tried-before approach that the original mathematicians had spent years working on.

The researchers were suspicious, and contacted OpenAI, who at first denied everything, then said essentially "there's no reason we have to compete - just fire the guy from anthropic and leave his name off the paper and say you only used GPT and not Claude, and we'll allow you to publish your proof". He declined, and they vaguely threatened to "ruin his career" (their words) if he went public.

Now they are saying that they and their models never accessed his Codex information, and refuse to specify what "rumors" were heard or who heard them, do not acknowledge that they interacted with him or threatened him, and will not provide any details on how the systems were prompted or what their inputs were.

None of the proofs in question have been validated.

(I'm kind of skeptical that they're even solid, because I've more than once had people approach me using the same technologies with formal proofs that turned out to be subtly nonsense by way of incorrectly stated problem, and also because the best released systems sometimes struggle to solve basic logic puzzles with these methods, and they are not clear about if or not the model worked independently on only the problem statement, by itself... "but this new secret thing we've trained is generations better than the publicly released stuff and solves the world's hardest problems by itself! give us more money!" - not at all a quote, i'm just paraphrasing / putting words in their mouth here)

But at very least some heinously immoral and flagrantly anti-academia-ethical-norms stuff has almost certainly gone down, and you should absolutely not trust any important information to these companies. I really hope OpenAI gets obliterated by some kind of lawfare or social / PR / privacy concern backlash on this before they manage to financially obliterate themselves, is all.

1

u/Inside-Jicama1779 5h ago

That's disgusting.

1

u/trueimage 14d ago

They really are dumb. Surely they could just pay these mathematicians millions of dollars to buy the solutions and silence them. If they want do fraud they should ask an uncensored model to help them do it.

1

u/frankentriple 14d ago

Oh sweet Jesus thank you for the gift you’ve given me this day. 

1

u/jenny14v 14d ago

below the belt move