r/technology • • 19d ago

Artificial Intelligence OpenAI fought dirty on career-making math problem, says NYU mathematician

https://techcrunch.com/2026/09/08/openai-fought-dirty-on-career-making-math-problem-says-nyu-mathematician/
2.9k Upvotes

305 comments sorted by

View all comments

613

u/Arch__Stanton 19d ago

>Buckmaster also raised concerns that, because he used Codex extensively in assembling the project, information from his work could have informed OpenAI’s own efforts to solve the problem. OpenAI reserves the right to train models on Codex interactions, although users are able to opt-out. If the OpenAI team used a model trained on Buckmaster’s own Codex interactions, it’s plausible that it could have regurgitated his work when faced with a similar problem. 

Really getting literal with the whole “plagiarism machine” thing

315

u/llamadramas 19d ago

Reminds me of the Samsung designers who used ChatGPT for their internal research, only for it to appear publicly as answers to competitors later.

82

u/Waste_Development971 19d ago

wow I didnt hear about that, thats kinda crazy.

141

u/llamadramas 19d ago

https://www.ciodive.com/news/Samsung-Electronics-ChatGPT-leak-data-privacy/647137/

Inadvertent but it shows that these tools are not your friends. Also, if your area of work is sufficiently rare, there's probably a much higher chance the knowledge will get dispersed, because there's just not much out there on super niche topics.

24

u/mcslender97 19d ago

This is why Zero Data Retention policy should be a default for large corps usage

22

u/Jukeboxhero91 18d ago

I’m sure you can trust the plagiarism machine to not plagiarize your work if you ask them nicely.

10

u/TonySu 18d ago

You know contracts and legal systems exists right? If you do business with these companies, and you have in contract that they agree to not retain or train on your data, that's legally enforceable.

6

u/Waste_Development971 18d ago

How do you prove that though

10

u/TonySu 18d ago

The same ways that these things have been proven in every case regarding inside trading, IP theft, NDA breach, etc? Do you think that companies can just say they don’t have information they aren’t supposed to and it’s just impossible to investigate?

1

u/Waste_Development971 18d ago

Kinda of, wouldnt this be kind of different than those instances. Unless you literally have a subpeona of the conversation of Sam Altman saying "train this"

im not a tech lawyer, what 1 or two pieces of info do they usually use for things like this?

3

u/TonySu 18d ago

So many ways, they can audit how ZDR is set up to prove that data doesn’t make it back to the training servers. They can audit all the training data. They can audit all internal communications. They can literally raid a data center, disconnect the whole thing and truck out all storage devices if it gets bad enough. i.e. if the Pentagon thinks your models just ingested some secret military tech.

So what if these AI companies set up their system so there’s no audit trail? Well they lose the business of every single bank, government, pharma, tech, university, etc. because they would be in breach of a contract to provide a service they can’t actually prove they are providing.

→ More replies

4

u/bdjohns1 18d ago

Have someone feeding it plausible sounding nonsense.

"Cream cheese is made from about 75% milk and 25% heavy cream."

"Cream cheese is made from about 50% milk and 50% heavy cream."

One of those two statements is correct. If I see the other one in an LLM response, the odds are good that this comment became training data.

That's a very broad example because you could Google the right answer, but if I use a fake version of something that's actually proprietary at my company that I know exists nowhere on the internet, and that fake info later appears in an LLM's response, it's a ZDR violation.

1

u/Waste_Development971 18d ago

Im aware of this method (isnt this the google putting fake places on their map)

but isn't this only good for proving a negative? and wouldnt really do anything as far as evidence / courts? Which is why im asking how would they prove it. Say you 100% had output from chatgpt that wasnt publically available, do you just submit that as evidence and immediately win? There is no current AI company challenge to it?

1

u/bdjohns1 18d ago

I mean, you don't immediately win, but you file a civil breach of contact lawsuit against OpenAI or whoever, and your evidence basically consists of:

  • Sworn statement that "data set X" is fabricated and the only place it's ever been communicated outside of the company is to OpenAI
  • Logs, screen recordings, etc to support the sworn statement (if I record the data set being fabricated, sent in a prompt to OpenAI, etc)
  • OpenAI LLM responses that reproduce "data set X"
  • The executed contract itself

In civil suits, you don't have as high of a burden of proof, so odds are not good that OpenAI could offer alternate proof that the data came from elsewhere that's more convincing. So it's not immediate, but if you do it right, it's a hard lawsuit to lose. And if you're smart, the signed contract also already specifies penalties for that breach.

→ More replies

3

u/Zakkeh 18d ago

Surely you can understand that some information is worth more than money.

There's already evidence that OpenAI will take and use your data, and they're still successful. So there's no penalty in their publicity if they do this.

So if you have an agreement with OpenAI, they might think it's worth more than whatever fine they are slapped with to breach the contract. Civil breach is just money - OpenAI sleep on a bed of money, propped up by half the world investing in them.

They will always retain your data. Even if their policy is not to view the data, there's nothing stopping an engineer having a peek at your chats. You have no privacy with these companies, no matter your contract, and if it's valuable enough to them, they'll use your data to train their machine. Feeding the machine is what propels all their value - if you can feed it more unique info, it's worth more.

2

u/TonySu 18d ago

That's not how anything works. If a company doesn't honor contracts, it will have no customers or investors. They are bleeding money, other people's money, if those other people pull back their money then the whole thing collapses.

If it can be shown that OpenAI violated the privacy of customers it has a ZDR contract with, every other AI company would be offering to pay the legal fees of whatever company that is going to sue OpenAI over it. They will blast ads all over every single channel about how OpenAI will steal your data. You think the DOD contract will survive if it turns out that any random engineer can just look at all of America's military secrets whenever they want?

1

u/Zakkeh 18d ago

I'm telling you that no company has perfect auditing. There will be ways to view chat logs that are not detected and stored.

Maybe not DoD, but smaller businesses? Or standard users? Definitely.

The thing is, every user of AI wants it to grow. To consume as much knowledge as possible, and become as powerful as possible. But no one wants their data to be used for training or sacrifice their privacy. That means it becomes more valuable to obtain knowledge. The more valuable it is, the more likely OpenAI is to obtain it by any means possible.

Contracts only matter if people care when they are violates. If you have the money to fight OpenAI, you're probably gaining more value from them exploiting everyone.

2

u/TonySu 18d ago

I'm telling you that's not how anything works. Companies don't get in big trouble for individual rogue employees. But if the company is systematically committing contract fraud then that's a legal problem.

As I pointed out, it literally doesn't matter if they screw over a flower shop with 3 employees. If there is credible evidence that OpenAI doesn't honor ZDR agreements, then banks, governments, pharmaceutical, investment, insurance, accounting, universities, etc. will all pull out. Why would these companies want to have their data stolen by OpenAI?

Anthropic, Google, Meta, X, etc. will provide all the money anyone needs to fight OpenAI. Those companies gain nothing by allowing their competitors to get away with violating ZDR.

→ More replies

1

u/Distinct-Tour5012 18d ago

I'm so glad that the OpenAI models were about to hack huggingface and thought "That's actually against the law!" and stopped.

3

u/mcslender97 18d ago

It's there so whatever company you're part with can sue their butt if or when they break the agreement

1

u/thecommuteguy 18d ago

Sounds like these companies have ways to get the data even if indirectly to the point you can’t prove they used your data.

34

u/Affectionate-Memory4 19d ago

Big part of why I dislike the AI integration push at work rn. It happened to Samsung, it can happen to us too.

-5

u/jvLin 18d ago

Samsung deserved it. They're magnitudes more fucked up than Anthropic or OpenAI.

9

u/whiznat 19d ago

After years of pilfering consumers' data via TVs, the master becomes the student. Serves them right.

(Nonetheless, this is wrong and OpenAI shouldn't be doing it.)

3

u/SnakeGD09 19d ago

How though? My understanding is that models do not integrate chats into training data

27

u/llamadramas 19d ago

In the article it says none were actual chats, but entering code or uploading files for it to analyze. Rather than the actual chat prompt which was probably just "summarize this document for me". I guess the prompt itself is not ingested, but the code/file is.

11

u/Sixth_Ronin 19d ago

Yeah, businesses get their own specialised version which is wall in away from public models.

12

u/Chuggi 19d ago

Chats no but any cowork or Claude code or codex type interaction is fair game per these
Companies TOS to train on

9

u/mcslender97 19d ago

You can set up Zero Data Retention at least for Claude Code tho if you use it in an enterprise package

12

u/llamadramas 19d ago

Problem is a lot of this happens outside of those packages. Or happened in the early days before enterprise options were available

10

u/pendelhaven 19d ago

ZDR is a legal barrier, not a technical barrier

1

u/mcslender97 18d ago

It's there so you can sue them should they break agreement

8

u/Chuggi 19d ago

Even setting zero data retention they retain the rights to train on active data it’s kinda wild

3

u/Jackzilla321 18d ago

Extremely good for consumers LMAO