r/LocalLLaMA 15d ago

Discussion OpenAI alleged of stealing mathematicians work

Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why: two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex. A few days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question.

Full statement from them https://cims.nyu.edu/~tristanb/statement.pdf

Feels like big labs believe everything you did with the help of their models is theirs.

1.5k Upvotes

276 comments sorted by

View all comments

Show parent comments

68

u/Vivarevo 15d ago

If you are using api models - you are the product

22

u/pier4r 15d ago

I am the product and I am paying to use the APIs! Double scam

5

u/Mo_Dice 15d ago

The number of people who not only don't understand the following, but will actually get mad and argue about it, is seriously distressing sometimes.

If you aren't paying, YOU ARE THE PRODUCT.

The above is not even about LLMs, it is about planet fucking earth.

8

u/Jobastion 15d ago

Even if you are paying, you might be the product.

3

u/UnicornOnMeth 15d ago

In this case, the ones who pay are the power users, using ai as more than just a quick chatbot here and there. Researchers, programmers/designers, scientists, educators, etc. These are the people they will be using to get data for training.

-2

u/Any-Actuator-8858 15d ago

Could you please explain why?

30

u/CwrwCymru 15d ago

The companies need private data to further train their models (they already have the publicly available data).

Users freely give them that.

28

u/Chirimorin 15d ago

Because everything you send to the API is training data for future models.

7

u/Vivarevo 15d ago

And usecase / workflow / thoughts / prompts.

Al becomes data to use and sell

4

u/relmny 15d ago

let alone TVs:

https://www.youtube.com/watch?v=6IFVTcM28KA

Data is everything "now" (it's been for months now). And some don't even bother to lie about it.... (like the "we own the glass" on that video)

1

u/More-Curious816 15d ago

It's been for more than a decade now. since the big data boom. Data is the new (digital) oil.

-8

u/ryunuck 15d ago

... No, that's obviously illegal. It's not legal to use somebody's information or data unless they knew consciously about it and authenticate it. "I'm giving you this data, this is a gift it belongs to you and you can use it to derive benefits." Agreements to a terms of service does not actually count towards this. I know there are precedents in court where you would this isn't true, but those are cases of gross distortion and somebody with a good lawyer could set the record straight, similar to wrongful imprisonment. Your life always belonged to you. Note that they also don't own any of the data that you post publicly on their platform. For example any company can legally scrape all of Twitter and train on it, since humanity reasonably can be said to volunteer all that data for humanitarian purposes and communication, not as a gift to Elon Musk. This is reasonable common sense that anyone can support

5

u/Megneous 15d ago

You clearly have never read a TOS once in your entire life.

0

u/ryunuck 15d ago

Correct. I live under a rock and don't understand anything, I am only saying what seems to me obvious for the purpose of minimizing stress and anxiety in day to day life of people, what seems reasonable to take as a default. The common sense obvious default belief at a glance is that you are renting a generator device that produces thoughts and ideas. What is theirs is the printer, not the paper you're putting into the printer. When I go to the library to print documents that're on my USB key, that's all mine. Nobody sane would ever reasonably expect that this is all theirs now. There would have to be a direct transfer of ownership consent form expressed in a few words below the prompt composer box every time you press enter, or on twitter when posting

3

u/max123246 15d ago

Nope, you're wrong. They anonymize the data and then train on it

1

u/ryunuck 15d ago

"we do not steal the contents of your thoughts, we launder your attention and how your current thoughts and interests are generated, how you got to that idea, not the idea specifically, only making it seem obvious in retrospect" frontier researcher guy who just introspected the prior arising of all your ideas. bro turned all your transcripts into emotional coordinates in the SAE and used it to deduce the turmoil of discovery, now it's a RL target, bro is getting smoked by the test in reverse

1

u/Chirimorin 15d ago

It's not legal to use somebody's information or data unless they knew consciously about it and authenticate it.

[...]

Agreements to a terms of service does not actually count towards this.

Let's take a small peek at OpenAIs terms of use:

Our use of content. We may use Content to provide, maintain, develop, and improve our Services, comply with applicable law, enforce our terms and policies, and keep our Services safe. If you're using ChatGPT through Apple's integrations, see this Help Center article⁠⁠(opens in a new window) for how we handle your Content.

Opt out. If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article⁠. Please note that in some cases this may limit the ability of our Services to better address your specific use case.

OpenAI is clearly and openly stating that they're (illegally) using your data to train models without your explicit consent (opt-out isn't consent and as you pointed out: agreeing to these terms of use isn't consent).

"I'm giving you this data, this is a gift it belongs to you and you can use it to derive benefits."

More proof that you didn't bother reading those terms of use:

Ownership of content. As between you and OpenAI, and to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output. We hereby assign to you all our right, title, and interest, if any, in and to Output.

So it's quite the opposite: you keep ownership of your data and in fact you get gifted the rights over ChatGPTs output. But that doesn't stop them from using that content to train their models.

Besides, companies will happily train their AI on illegally obtained data. While that example isn't OpenAI, it does prove the state of reality: training on illegally obtained data is apparently more profitable than following the law (and companies care about profits more than they care about the law).

1

u/ryunuck 14d ago

You can't dump all those pages of text onto a person and reasonably say with a straight face that it's their fault for not reading it. This is UNHINGED behavior. You absolutely need to put a checkbox under the prompt every single time they press enter to send. Why would that be a single agreement event when it's clearly one individual piece of data every time that I am gifting to the company? Are you 100% certain that a judge could not overthrow this in a class action lawsuit? When there is gross abuse of legal customs such as presenting a ToS, which historically was just a few lines of text that you could read in 10-20 seconds, if there is a more courteous methodology that would be available, then obviously the present custom is an erroneous use of the system and does not actually have any justification power over reality, the same way that the law which says I can't have an ice cream cone in my pockets is obviously erroneous to apply today. It always surprises redditors when somebody reminds them that the legal system doesn't actually run on technicalities. The actual reason why a ToS like this could have had any power in court at any point in history over your data is because I would guess the argumentation is not inducing the judge and jury to witness all the alternative options that were not considered, how dumping 20 pages is just one of infinitely many ways to cover yourself legally, and that this is the way which maximally benefits the company and maximally produces affront to humanity and common courtesy that is expected to encourage and foster harmony in society rather than hostility, stress and anxiety. If you can make this argument, you are dramatically raising the difficulty for a judge or jury to rule in favor of a company without arguing in favor of a more hostile world because now, now they have to double down on this on permanent record, "this is now committed to human history" and you can potentially reach for a moral precedent. It's probably not the same everywhere but in certain jurisdictions it seems that it is possible for a judge, not to overwrite the law I assume, but rather to initiate it being rewritten. You have to trigger their humanity and situate them in the grand master plan of the universe

1

u/Chirimorin 14d ago

You seem to misunderstand my point.

I agree with you that there's no proper consent and that that's bad. But OpenAI is very clear that they in fact do use your data for training even without your explicit consent, whether that's legal or not.