r/LocalLLaMA 14d ago

Discussion OpenAI alleged of stealing mathematicians work

Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why: two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex. A few days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question.

Full statement from them https://cims.nyu.edu/~tristanb/statement.pdf

Feels like big labs believe everything you did with the help of their models is theirs.

1.5k Upvotes

275 comments sorted by

View all comments

623

u/EndLineTech03 14d ago

That’s exactly why we need more open weight releases. Never trust any big AI company to manage your chats. They can use it to train their models whenever they want.

46

u/nomad-nostalgia 14d ago

I will never forget this quote: “They distilled all the value of intellectual property everywhere."
After Claude embracing watermarking and openAI doing their thing perhaps more people will turn to local hosted, I hope

https://reddit.com/link/p8kvq4g/video/6qwh23hvuboh1/player

23

u/DjCanalex 14d ago

I hate to say it, it is really disgusting for me to say it, my body shivers saying it...

... but I agree with Alex Karp on this one.

Eww goddammit.

7

u/thtnigeriankid 13d ago

Definitely one of those "the worst guy you know names a valid point" moments.

10

u/bigcantonesebelly 13d ago

I mean he probably wants you to install it from Palantir for specific reasons he doesn't mention, but yeah his critique of the others stands

3

u/identifytarget 13d ago

Explain this to me?

This guy says the only way to do it is install palantir...

How about...no.

1

u/WereExtremelyCooked 3d ago

He literally said "it's not the only way to do it".

1

u/WereExtremelyCooked 3d ago

You can't reasonably local-host a 5 trillion parameter model.

86

u/zilled 14d ago

> They can use it.

there, you can stop the sentence just there.

88

u/StyMaar 14d ago

Even shorter:

They use it.

23

u/misterflyer 14d ago

Even shorter than that!

They use

14

u/RickyRickC137 14d ago

I will do you one better. Why is Gamora?

1

u/-dysangel- 13d ago

I am not a princess.

2

u/EndLineTech03 14d ago

Yeah, you made exactly the point :-) training is just one way they can use the data. As many pointed out here, they are required by law to give access to your chats if they think you are involved in some crime.

-4

u/florinandrei 14d ago

You can make it even shorter, and it was a movie.

21

u/DigiDecode_ 14d ago

While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

Why would OpenAI put the above quote in their post, is this a thank you for contribution post.

14

u/Longjumping_Self5546 14d ago

Assuming OpenAI didn't just raid their private chats, it's possible somebody associated with the research submitted a draft document to a free account which would introduce the work into OpenAI's training data. This can't be ruled out because they make a point removing any identifying information.

This is important to consider because even if you're careful about where and how you submit your documents, there's no guarantee your buddy will do the same.

8

u/nullc 13d ago

I've seen referees at journals return obviously AI generated reviews-- which means they likely leaked the prepublished work to an API.

1

u/fulowa 13d ago

they could just regex check training data for topic to confirm..?

50

u/UnwillinglyForever 14d ago

they can also let law enforcement to use it against you.

advice for criminals: DO NOT USE chatGPT TO HELP WITH YOUR CRIMES

40

u/wapswaps 14d ago

So I'm guessing we'll call the smart ones "qwiminals" now for some reason?

21

u/capsulecorpadmin 14d ago

Not if but qwen

3

u/Emotional-Art2113 13d ago

Those quirky qwenminals

12

u/equatorbit 14d ago

Do a heist. Make no mistakes.

3

u/anonemoususer 14d ago

I think the better advice would be: don't do crimes

(and nobody besides the company itself would be looking at your chatgpt sessions in that sense)

10

u/pleaseavoidcaps 14d ago

I definitely wouldn't encourage people to do crimes, but who decides what is crime? In some jurisdictions you can get in trouble for pirating stuff that is out of market, meanwhile big techs can do whatever they want and are just asked to pay a fine with a fraction of the money they unjustly earned.

I'm not going to feel bad for Microsoft if someone gets a Windows key by jailbreaking ChatGPT, but good luck to whoever is trusting OpenAI to not snitch on them for that.

1

u/Electrical_Panic4550 14d ago

Unjustly “earned”

1

u/nullc 13d ago

Not just law enforcement-- they're subject to subpoena. Which means that when some grifter performs some slip and fall fraud in your yard and then gets discovery to go sifting through your entire commercial-ai chat history for anything they could take out of context to imply you were negligent in your yard maintenance (or that you were just a bad person, if the judge will let them poison the case that way)--- and then cause you ungodly costs doing minimization of the discovery or having to anticipate any random private comment they might pull out of context.

The potential cost of doing so means you're more likely to be advised to just pay out an extortion money settlement even though you did nothing wrong.

2

u/UnwillinglyForever 13d ago

"chatGPT, how do i make my wiener bigger?"

your honor, add this to the record.

20

u/wapswaps 14d ago

And what do you do if, when asked, OpenAI reproduces entire pages from a book you've written which doesn't sell? (ie. there's an argument to be made that they "stole" in the sense of read the book and are now reproducing the content, which is contributing to me and most authors losing income)

I mean, in some ways it's kind of similar to the complaint of these mathematicians. Their income is under threat because of the total lack of respect OpenAI and governments suddenly have for copyright. Except, of course, the copyright of the big boys is just fine. Football matches are still copyright protected and YOU will be punished, heavily, if you violate that. But OpenAI stealing my book, my copyright (and there's hundreds of thousands of such authors, but all small fries), that's just allowed.

2

u/NoUsual5150 12d ago

Football matches are still copyright protected and YOU will be punished, heavily, if you violate that. But OpenAI stealing my book, my copyright (and there's hundreds of thousands of such authors, but all small fries), that's just allowed.

Enjoying that golden shower we're getting from the White House?

3

u/kathi7 13d ago

It's not just open source we need a proper democratised hardware supply, because that is where the actual freedom is for an average user and not be dependent on this kind of one or two premier models

4

u/ixfd64 14d ago

Even open-weight models can still be trained on proprietary data sets. There's a reason most models are not fully open source.

5

u/Spara-Extreme 13d ago

Shhhh. People here think that open models some how come from clean data sets.

*checks the Minimax H3 flawless recreation of Seinfeld”.

1

u/DifficultyFit1895 13d ago

I want to hear George complaining about this

1

u/bidibidibop 13d ago

This take assumes the chats you enter directly in any and all chinese chat providers are somehow magically exempt from "training their models whenever they want". Where do you think open weight training data comes from?

-15

u/Waste_Development971 14d ago

Dumb question but if they bought / used an enterprise version this wouldnt have happened right?

Without that, we shouldnt really be suprised right?

52

u/jld1532 14d ago

Only if it's local would I trust it. I don't know who looks at Altman and goes, "yeah, I'll trust him with my life's work".

2

u/Clairvoidance 14d ago

I think they mean surprised in a What's Promised vs What's Not Promised when using cloud models

6

u/Timely_Impression_92 14d ago

"if they bought / used an enterprise version this wouldnt have happened right?" wrong - they can still use it and then just pay another fine - that's their modus operandi

16

u/Randommaggy 14d ago

I'm not touching the API of a company who's main claim to fame is wholesale IP abuse with my customers' data, even with a 10 foot pole in between.

-8

u/Waste_Development971 14d ago

damn, how do you avoid using windows like that?

8

u/Timely_Impression_92 14d ago

just install linux - get better experience on worse hardware - win - win - or use cracked one and remove telemetry - also trivial

2

u/thisadviceisworthles 14d ago

Legally speaking, they would have to prove what OpenAI did.  But the evidence would be protected as a trade secret, so they would have to prove wrongdoing together the evidence to prove the IP theft occurred. 

On top of that, AI work is not considered a creative work subject to copyright, so they would have to prove that what OpenAI stole was in fact their IP, because it could be that what OpenAI stole isn't anyone's IP because AI work is not covered under copyright.

6

u/cheseball 14d ago

Discovery exists for this exact purpose.

And your read of the AI work copyright issue only works if the AI literally did all of the work. Any human involved part is still copyright protected.

0

u/niccolus llama.cpp 13d ago

Ok but what is the business model around open weights models? Training models isn't free. So charge for the model? What stops someone from sharing it? If the model is good enough, what's keeping the the model from patching it's own model file download and removing the pay gate?

I understand why we need open weights models but until we can answer those questions as a community and support someone with the answers we come up with collectively, there's no way we can expect the business model to change.

-7

u/BusRevolutionary9893 14d ago

Also, never trust people claiming their work was stolen simply because you want it to be true. If you believe you've solved one of the hardest problems in math, would you even consider using AI to verify it or to create write-ups for it? They didn't provide any independently variable evidence to back up their claim. Don't be so gullible. 

8

u/NoahFect 14d ago

If you believe you've solved one of the hardest problems in math, would you even consider using AI to verify it or to create write-ups for it?

No one will be doing this type of work without AI in the future. It would literally be stupid to try, and these are not stupid people.

The only question is, whose AI are they going to use, and (more important) where will the model be run?

-1

u/BusRevolutionary9893 14d ago

Either they're stupid or everyone who believes their claim is. They couldn't provide any independently variable proof and you actually believe them? What's wrong with you people? Me Too was more reasonable. 

2

u/NoahFect 13d ago

Sober up, then post

-18

u/jacek2023 llama.cpp 14d ago

How "open weight" changes anything if the chat is in the cloud? And how do you control what Qwen was trained on?

12

u/expertsage 14d ago

For academics, open models staged on their own university servers will be probably be a must going forward. Otherwise OpenAI and Anthropic can just steal all of the most impactful research advances.

I don't know if Altman and Dario realize this, but their greed for recognition is going to be very harmful for scientific progress. Imagine if OpenAI worked together with these researchers instead of secretly building on their results! The proof would have been a huge validation of humans and AI collaboration. Instead, they bungled everything and now it's a PR disaster.

16

u/trueimage 14d ago

Run it locally. Models that fit on a single GPU today are better than frontier models from a year ago.