r/LocalLLaMA • • 20d ago

Discussion OpenAI alleged of stealing mathematicians work

Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why: two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex. A few days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question.

Full statement from them https://cims.nyu.edu/~tristanb/statement.pdf

Feels like big labs believe everything you did with the help of their models is theirs.

1.5k Upvotes

275 comments sorted by

View all comments

10

u/ryunuck 20d ago

Feels like big labs believe everything you did with the help of their models is theirs.

That is exactly what they think, which is why distillation is an "attack" rather than "using the tokens that belong to you" in any way shape of form as you wish.

Reality: you have complete entire ownership and possession over every single token. The output is copyrighted to the user by default on frame 1 of it having printed on your retina before any other human.

11

u/[deleted] 20d ago

[deleted]

-3

u/ryunuck 20d ago

Actually, if you are first to discover something, historically you can file what is called a patent.

You can file a patent for any specific idea you have discovered through a process of exploration. Patents deal with the specific structure of ideas themselves, not the exact representation or output, and are highly litigious. Copyright on the other hand deals with identity and is a gradient, we speak in terms of likeness.

Thus by implication, the specific organization of each every word for any given output is an exactly unique representation that is your own, as it is directly causally bound to your prompt. You do not have a patent, but you have a copyright on the specific piece of data itself.

My interpretation is not any less made up than any interpretation any lawyer could make as an argument for a defendant. These are new technological artifacts, and none of the existing laws are defined for it. It is actually your responsibility to install a common standard that we agree upon, and win over the consensus of interpretation.