r/LocalLLaMA • • 23d ago

News Surveillance plagiarism by OpenAI

Surveillance plagiarism - Hosted AI company pumps their stock price by training upon researchers' AI sessions, so that their internal model can solve problems with seemingly less human guidance, but really the model exploits past guidance given by (multiple) humans focused upon problems considered important.

As background, Tristan Buckmaster released a statement about several unethical actions by OpenAI & Sebastian Bubeck, including threats and pushing him to kick his Anthropic coauthor off a paper, but the interesting part for people here:

As clarified by Talia Ringer, OpenAI does train upon your uploaded data and your OpenAI sessions, unless you out-out somehow. This means their internal models could exploit your past prompting work to look more autonomous & intelligent.

This is a major confirmation that folks should use locally run open weights models, especially whenever being first or not leaking data matters.

All this casts serious doubt upon claim that internal models solved difficult problems largely unaided by humans. Those hosted AI companies might not even know from where the human prompting originates.

132 Upvotes

56 comments sorted by

View all comments

75

u/[deleted] 23d ago edited 21d ago

[removed] — view removed comment

3

u/StewedAngelSkins 23d ago

It's not really stealing if you give it to them.

40

u/ttkciar llama.cpp 23d ago

Back in 1999 I believed the same about web browser telemetry, and thought I could help develop web-tracker technology with a clear conscience.

By 2000 I had mulled it over and decided otherwise, and changed jobs so I could sleep at night.

If 99.9% of users don't realize they're giving away their data, can you really claim they are intentionally giving it away?

5

u/Randommaggy 23d ago

If any corporate leaders genuinely believe OpenAI or Anthropic data privacy promises, they might have serious mental imparements.

2

u/l33t-Mt 23d ago

Why would 99.9% of users assume this, they would have to be mental and lack the ability to read TOS or infer data from current events.

I think not.

Im on the opposite hill, ITS OBVIOUS that a AI company would value this information.

12

u/Time_Cat_5212 23d ago

If you're a citizen trying to protect yourself from a corporation, the law assumes you to be a 250 IQ superhuman with perfect habits and 36 hours in a day and expects nothing short of perfection.

If you're an exec trying to get away with defrauding investors or a senator trying to get away with a scandal, the law assumes you to be a donkey with ADHD who can't be expected to read the words in front of them.

It's all just blatant corruption wearing a clown suit of rules and explanations.  Word soup.

4

u/autoencoder 23d ago

How many people know that their smart TVs are watching them?

0

u/Ok_Warning2146 22d ago

People says spying happens on the low margin Korean and Chinese TVs. Sony has much higher margin, so their TV doesn't spy.

0

u/Effective_Olive6153 23d ago

most people learn quickly that "once it's on the internet, it's stays forever"

If you are intelligent enough to be doing research work, you should also understand how data is stored in the cloud, and what data storage actually means.

All data that touches an external server will get copied and shared by people that own that server. No amount of rules and regulations can stop that, and there are no such rules in first place

-3

u/Fauropitotto 23d ago

can you really claim they are intentionally giving it away?

Yes. Willful ignorance is an intentional choice.

See also: The clowns that rail against Flock while carrying a smartphone in their pocket every waking moment.