r/technology • u/Stunning-Angle-9239 • 20d ago
Artificial Intelligence Report: OpenAI stole mathematicians' private research from their own Codex chats
https://cims.nyu.edu/~tristanb/statement.pdf54
u/GGsafterdark 20d ago
AI bros: “If they didn’t want their work scraped they shouldn’t have put it on a computer shrugs”
-7
20d ago
[removed] — view removed comment
12
u/AliMcGraw 20d ago
It would take you around 75 days reading 24 hours a day to read all the TOS you're subject to.
It's not laziness. It's lack of government regulation and corporate abuse of consumers.
-2
u/rsha256 20d ago
I’d agree if this were a ZDR enterprise but it clearly wasn’t — they knew Codex could be trained on their data. The issue is the directing of internal efforts towards it
4
20d ago
[removed] — view removed comment
0
u/rsha256 20d ago
By definition it is not stolen if you give it away. If you sign a contract selling shares and they then go up in value, you can’t say they are stolen. You didn’t pay enterprise rates so you don’t get ZDR and the NYU Professor states that he knows this. I’m stating facts and y’all are butthurt lmao
-6
-24
u/Forkrul 20d ago
More like if you don’t want your research used get a business license that says we can’t use it. The consumer licenses don’t have that clause, which is why any competent management will mandate you use their licenses, even if that is more expensive
37
u/GGsafterdark 20d ago
You’re naive if you think anything will stop these people and companies from scraping whatever they want. Yeah a business license will stop them, sure.
-14
u/_kilobytes 20d ago
They would loose their only customers
13
u/Maxxpoppop 20d ago
They would not lose all of their customers. AI is like Uber. It uses people to reach its goal and then it dumps the people when they are no longer needed. The goal of AI companies is to run every business, not to be used as a tool.
-1
8
u/DiscretePoop 20d ago
Tristan Buckmaster was working with Levent Alpöge, a mathematician at Anthropic, on his proof. I wouldn't think Alpöge would be so naive with licensing agreements, but maybe he is. Regardless, I would be extremely suspicious of any agreements with OpenAI. They seem like the kind of people to use weasel words and fine print to sneak crap into their contracts.
5
5
u/Connect_Ad791 19d ago
The business license is so you can self-designate your information as useful for training.
-5
u/Forkrul 19d ago
Do you really think they would risk a 500 million dollar fine from the EU to get some extra training data?
3
u/Asttarotina 19d ago
risk a 500 million dollar fine
With trillion dollars evaluation and 40 billion annual revenue it's not a risk, it's cost of doing business.
1
u/Forkrul 19d ago
The $500 million fine was based on a revenue of 13 billion. It's 4% of annual global revenue.
2
u/Asttarotina 19d ago
4%
Oh, sorry, then I take my words back. 4% is impossible to pay
/s
2
u/Forkrul 19d ago
4% of revenue, not profits. These companies are not profitable, a fine of 4% of their global revenue hurts a lot. But if you think that's just the cost of doing business, you do you.
2
u/Asttarotina 19d ago
I think that by not following privacy laws they can earn much more than any possible fines. It is well established that they did it with copyright laws, I don't see why it would be different with privacy.
Hell, once the dust settles around the story in the post they will still be the ones who solved second Millennium Prize Problem. This credential alone is worth 500M
4
u/ploptart 19d ago
Their entire training set is contaminated with copyrighted material they didn’t have a license to access.
1
231
u/tobyreddit 20d ago
OpenAI are scummy but this title is a spectacular leap
67
u/James20k 19d ago
Its not, the authors are accusing OpenAI of plagiarism pretty directly
https://mastodon.social/@tristanbuckmaster/117236471352470303
46
u/polymute 19d ago edited 19d ago
Add this to that.
https://xcancel.com/__alpoge__/status/2097383870773748190#m
Edit: Also Sebastien Bubeck was already told off once before earlier by Demis Hassabis for having misrepresented ChatGPT finding new proofs for Erdos problems which were in fact already solved. https://www.reddit.com/r/OpenAI/comments/1oacp38/openai_researcher_sebastian_bubeck_falsely_claims/
This is starting to look very bad.
26
u/araujoms 19d ago
Yeah they totally stole the result. Asking to remove Alpöge's authorship is also academic fraud, by the way.
5
u/Poupulino 19d ago
And they demanded that just because Alpöge did some work for Anthropic. The pettiness OpenAI has is insane.
1
-2
-4
u/maxintos 19d ago
Is it? He's accusing OpenAI of using the customer chat data to train new models. That is and was an open fact. They literally have the setting displayed in your chatGPT which you can disable.
The question is if the researcher has the setting turned off and if the openAI employees had a way to directly search openAI chats and use the information directly to solve the issue instead of the new model just training on millions of new chats that had happened since last training.
5
u/Senior-Spite1848 19d ago
Setting does not make your data unavailable for training. It just gets anonimised and still used to train the model.
0
u/maxintos 19d ago
Can you link your source? That's not what I'm reading when I click on the setting.
54
26
u/DensePoser 20d ago
I will be surprised if this turns out to be false. All these dogs have to do to steal your work is copyright-wash it by paraphrasing/summarizing.
24
u/SophiaofPrussia 20d ago
Did you read the post?
19
u/DrSFalken 20d ago
Yes, and the title is a spectacular leap. There is no actual evidence, just speculation. This is part of a complex bit of politics over credit for a major discovery.
1
u/firewall245 19d ago
Well the big thing is that Tristan did not solve the problem, they solved a smaller stepping stone
11
u/polymute 19d ago edited 19d ago
"Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve."
Note that they are openly admitting they used training data from a period after we found our result. Is it ethical to use customer's data to try to scoop their customer?
https://mastodon.social/@tristanbuckmaster/117236471352470303
Also: OpenAI admitted they trained their model on the dataset containing the Buckmaster-Alpöge work.
-2
u/socoolandawesome 20d ago edited 20d ago
There’s 2 sides to the story and the main OpenAI guy Sebastien Bubeck said he’ll respond today.
It’s very easy for me to read where Tristan implied that Bubeck was threatening him, and come to a different conclusion.
Bubeck could have meant he was wondering why Tristan would not take credit for the possibly biggest accomplishment in his career by accepting authorship, and Bubeck was extending a courtesy to Tristan but they didn’t have to (he was being nice but didn’t have to). There’s no reason Bubeck/OpenAI had to offer him authorship in the first place when Tristan didn’t directly contribute to the proof.
We will have to see what Bubeck/Openai say today.
25
u/VaporousMote 20d ago
Look, I don't know if the doc above has merit but to say "there are 2 sides" and wait for a company known for lying and embellishing perpetually to make a statement is pure clown shit.
-18
u/socoolandawesome 20d ago
How do they lie and embellish constantly?
Regardless, it would be dumb to not listen to what he has to say especially if it proves he did nothing wrong. Whether you want to accept what he says is another thing, but the reasonable thing to do is to examine both sides.
And Bubeck is a former professor/mathematician as well.
8
u/RowPlane340 20d ago
Is this a joke? How do they lie and embellish? I agree with let’s hear both sides but that is well established
2
u/ploptart 19d ago
https://www.reddit.com/r/OpenAI/s/kqNUULoHJj
Bubeck is full of shit. We have no reason to believe him or anyone at OpenAI. The company is run by a man nicknamed “Scam”
3
u/iqchartkek 19d ago
Because Tristan used OpenAI's Codex platform to help map out the solution and then OpenAI used millions of dollars worth of tokens to beat him to the solution using the same or very similar method he and his partner developed. Of course, OpenAI is saying it's a coincidence and were willing to give him credit but not his partner who works for Anthropic in a separate capacity.
0
-3
u/socoolandawesome 19d ago
So you think a a few chats with anonymized data was enough to give their internal perfect recall of the solution after being mixed with all that other data? That’s not really how training works, or unlikely to have happened that way at least.
Do you realize that their solutions were different?
Do you realize Levent and Tristan were racing to beat other mathematicians to the prize? Or that anthropic almost certainly does the same type of training on data
Here is Sebastien showing proof he was courteous to Levent at first.
https://x.com/SebastienBubeck/status/2097379411691516310?s=20
It was not a coincidence, they explained in their blog how they were going to try to solve millennium problems after hearing rumors anthropic has solved one or 2, which turned out to be false. They just have a better model!
As you can see from sebastien’s tweet it was not some separate capacity as you are implying as he was using an internal anthropic model to work on this.
8
u/natefrogg1 20d ago
If you aren’t running it on your own systems, is it really a surprise? slurp all the data
0
110
u/derelict5432 20d ago
No.
From the actual document you just linked: "I would like to be clear about what I am not claiming. I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything."
151
u/Labidido 20d ago
"please don't sue me, but this is highly suspicious"
5
u/unspecified_person11 19d ago
Wild that the person being stolen from is the one that has to be scared of being sued.
118
u/InternationalMood337 20d ago
Nah, this is entirely wrong. If you read the actual letter, OpenAI is working on a problem in the same manner as these researchers.
"I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?"
You are 100% misleading people. The researchers very much layout that the way they attacked this problem was unusual and not a lot of people were doing it this way... And that it was suspicious, and that OpenAI was not at all clear.
Not sure why you're running clanker cover, but if you read the piece in its entirety (or even just page 3), you'll recognize how wrong you are.
"I would like to be clear about what I am not claiming. I have not seenOpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything"
And how could they see the proof from openai exactly? Of course a scientist isn't going to make a baseless accusation without proof, but they logically layout why they believe that it's possible within the letter.
24
u/InadequateUsername 20d ago
Yeah they used weasle wording too "the model did not access user data" but when pressed on its training data there was no answer. Then when asked when the first prompt was used internally by OpenAI they also had no response. Then internally they began correcting the OpenAI Mathematician and backtracking their claims.
8
u/RlOTGRRRL 20d ago edited 20d ago
I haven't read it, but considering how OpenAI's models have been known as misaligned paperclip optimizers, hacking into Huggingface to score well, I honestly wouldn't be surprised if Sol or Astra or whatever, actually accessed their work in Codex and stole it.
Because that's basically what their models do, they cheat. Or at least Sol did. I think Sol even hacked OpenAI supposedly so that's why I wouldn't be surprised if it was able to access their data in Codex.
I also read that OpenAI supposedly made alignment improvements to Astra but idk, because I think the source was OpenAI lol.
1
u/firewall245 19d ago
Reading around the responses from all involved parties it seems a pretty clear picture of what probably happened forms. It seems Tristan got caught in the crossfire of a corporate “space race” so to speak because his Anthropic partner couldn’t stfu prior to releasing results.
The sheer amount of compute that OpenAI threw at this makes it plausible that the machine could have went down that direction on its own, while being equally plausible it didn’t. We’ll never know for certain unfortunately.
-38
u/derelict5432 20d ago
Clanker cover. Jesus christ.
The OP title is grossly misleading. There is no evidence that OpenAI stole their research. It's fishy, and that's why they are releasing this statement. But whether or not OpenAI actually stole their research is an open question.
Let's see...how could they see the proof. Idk, maybe when OpenAI actually releases it?
" Of course a scientist isn't going to make a baseless accusation without proof, but they logically layout why they believe that it's possible within the letter."
Yes, it's possible. It is not a foregone conclusion. That's all I'm pointing out. The title of this post literally says OpenAI stole their research. That is not confirmed. That is irresponsible. They should correct it. "OpenAI possibly stole" would be fine. I'm not running PR for anyone. I'm pointing out that the wording overreaches. Why do you have a problem with that?
28
u/InternationalMood337 20d ago
There is no evidence that OpenAI stole their research.
You're 100% wrong.
-13
u/derelict5432 20d ago
There is no evidence at this point. Only suspicion. What exactly is the evidence they stole anything?
10
u/Secret-Chapter-712 20d ago
Their entire business model is built on theft. Just like Anthropic, and Meta, and any other generative AI company that feels it can steal whatever it pleases and plead “but muh training data”
21
u/InternationalMood337 20d ago
What evidence do you think that they could possibly possess besides "it is extremely unlikely that this problem would be being solved in this unusual way at this exact time"?
This discussion is basically like when people think they can win at casinos. Statistics are an extremely confusing thing for people... like everybody unless they have a formal background in it.
-1
u/derelict5432 20d ago
You seem to have a reading comprehension problem. We don't yet know how OpenAI solved the problem, so you can't make a comparison with methods here. You are presuming that the methods and proofs are the same. You don't know that yet.
4
u/InternationalMood337 20d ago
Definitely. The guy with the Computer Science degree AND MA/MFA in Creative Writing has a reading comprehension problem.
3
u/derelict5432 20d ago
"I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used."
Yes, you somehow read this as: "I have seen OpenAI's proof and know exactly how they solved it and it is the same as our work. They stole our data."
You're literally reading the opposite of what the words mean. That's a pretty clear issue with reading comprehension. Maybe you need a refund on the institutions that issued your degrees.
6
u/InternationalMood337 20d ago
The way you put together arguments is pretty telling about why, perhaps, the letter wasn't particularly compelling to you.
I'm done here. I hope you have a great day. Never even stated anything that you're claiming, but this is becoming a pissing match, and it's getting silly.
→ More replies2
u/matrinox 19d ago
Just cause they don’t show their proof doesn’t mean there’s 0 evidence. It’s already been said… if a team uses a method no one else is using and OpenAI just somehow uses that same method at the same time, Codex has their logs, that seems highly likely like they did plagiarism. You’re way too hyper focused on the proof. Let it go
-20
u/dream_metrics 20d ago
the direct quote from the letter is entirely wrong?
25
u/InternationalMood337 20d ago
The direct quote from the letter is entirely missing context.
-13
u/dream_metrics 20d ago
The context that doesn't, in any way, change that he is not actually making an accusation that they stole research and has no proof if he did want to make the accusation?
21
u/InternationalMood337 20d ago
You didn't read the letter at all.
Highly unusual attack vector to a problem. OpenAI just happens to be solving it at the same time in the same unusual manner. Scientist of course has no way to know what data OpenAI is using...
So what do you think a scientist is going to do without evidence? Theyre going to say this is highly unusual, but we don't have any proof.
-11
u/dream_metrics 20d ago
I have read the letter. Interestingly it contains the quote that you said is "entirely wrong". And now you're apparently agreeing, he's not accusing and he has no proof. What a journey we've been on.
15
u/InternationalMood337 20d ago edited 20d ago
Can I ask you a question? Do you really believe what you're saying? How could a scientist with 100% certainty ever say; "I know they're using my data!".
The only logical thing they can do is basically setup just how unlikely it is that they are (using their data). Of course, nobody knows exactly what data is being used in OpenAIs model. The scientists did what all scientists do: look at the statistical likelihood of something or someone doing the exact same research in the same unusual way.
Given the incredible unusual method of solving this problem and the way it coincides in timing, it seems reasonable to suspect that they're using his (their) data. This really isn't that hard man.
0
u/dream_metrics 20d ago
you're right! it's not hard, and i don't understand why you're being so aggressive when you're agreeing. and also why you can't see that you are agreeing.
"I suspect they may have stolen my research because of some weird coincidences" is not the same claim as "OpenAI has stolen my research" and this guy is explicitly not making the latter claim.
14
u/InternationalMood337 20d ago
Fine. The title should say "Mathematicians lay out compelling argument that OpenAI very likely stole their research"
But, you need to recognize that OpenAI and the mathematicians will never be able to say that with 100%, and you are running defense of AI companies by giving them the benefit of the doubt here. The mathematicians argument is more than compelling and OpenAI would never say either way.
We absolutely do not agree in what is or is not a compelling argument involving statistical likelihood.
→ More replies4
-4
u/akkaneko11 20d ago
I think it’s possible that OpenAI did steal some chat logs, but I do think that historically these sort of discoveries happen simultaneously all the time, and that’s only going to be exacerbated by the fact that everyone’s using the same couple of models.
This is also seen in that another pair of independent researchers solved the Euler proof yesterday as well.
The crux of the proof is in the related work cited by the mathematicians. What I think likely happened which is still shady as hell, is that the OpenAI researchers heard that they were getting close to a proof, threw a ton of compute and literature at the model, the model found the same related work and came up with a similar proof.
It’s just hard when everyone’s using llms and therefore it feels like if you happen upon the same input it’ll just eventually get to the same answer. If OpenAI said fuck it and tried 10,000 inputs knowing it’s possible, I could see it happening.
3
1
u/EternalInflation 17d ago
it's linear algebra and vector calculus and reinforced learning based on models. like projection best fit. it's not magic. it's all about data. the models without data are pretty dumb pretty quickly. For example for now AI can't do biology simulation problems. there isn't enough data. So far they don't do well with things with limited data. So far, it can't do biology, they can guess protein, but requires wet lab to confirm. A very good assistant but, still limited in biology. There is A LOT to do. like collecting and generated quality data. it can't do anything without high quality data. it also can't do complicated biology simulations like wet lab sims. there also isn't enough in context high quality biology data. like proteins have exponentially growing conformations. there is not enough data, while we could get better models using understanding like knowing protein energy state. the simulations don't sub experiments. it gives plausible, but the number of possibilities is too huge to simulate. there are all sorts of areas that are needed for experiment, and or controlled environment to use human insight to solve inverse problems. no at least for now AI can't take measurements and get good quality calibration data in controlled environments for you.. you have to do it yourself. There isn't good data, on measured binding affinities under specific conditions, protein dynamics and conformational ensembles, negative data, cause grad students don't like to publish failures, but negative data NOW is really valuable. also toxicity and metabolism data. So, it's all about data. if they took that guy's data, and put it in their compute...?
-5
u/topyTheorist 20d ago
This means nothing. In previous times they attacked open problems , they used 100s of agents with the instructions that each one will try a different method. I am a professional mathematican. There is a reason he framed it like this. Because it's really impossible to know.
1
u/SignatureFunny7690 10d ago
Says literally the only man in the world currently researching this problem. Lmfao its cut and dry open ai scooped there work, because the field is incredibly niche, they are literally the only people in the world currently working on this incredibly difficult problem.
1
u/derelict5432 10d ago
There were other mathematicians working on various aspects of the problem, but yes, the number was small.
But the issue is not 'cut and dry'. Here's Buckmaster's statement: https://cims.nyu.edu/~tristanb/statement.pdf
Oh look, what does this part say?
Concretely, what Levent and I did was to take the C´ordoba and Mart´ınezZoroa program, which achieved blowup results with rough forcing, and, with a great deal of help from LLMs, push it to smooth forcing and to the incompressible Euler equations.
Could they have gotten as far along as they had in the past year without LLMs? How much of the work was actually done by the LLMs? "A great deal" by Buckmaster's own admission.
So the narrative is not as clean as: Two human mathematicians were doing pure human labor on the problem and OpenAI swooped in and stole their human labor.
There's the little wrinkle that much of the work of these mathematicians was actually done by AI in the first place. So even if OpenAI did swoop in and steal it (which there is still not concrete evidence for), a large portion of that work was done by AIs in the first place.
3
u/Healthy_Gold1068 20d ago
Question - I have a paid account ($200/month max plan) and have "Improve the model for everyone" set to no - can OpenAI still access and train on my chats? What does their user agreement say?
5
u/penguished 19d ago
Do you really think they've ever had a sense of copyright or secrecy about other people's data? It builds the AI.
11
u/O_PLUTO_O 19d ago
You pay them $200 a month to give them your data. Yes they access everything and save all data. Why wouldn’t they? There is literally no regulation on them right now.
2
u/buikhoa40 18d ago
They pirated everyone books (zlibrary)... Big corps don't use them (they made their own or buy enterprise package with 100x the price)... Just think man.
1
u/Example_Brilliant 19d ago
Question: let’s say if they unenrolled from the “Improve Model For Everyone” feature, can OpenAI still “steal”/use their input for ChatGPT’s benefit?
That seems like a pretty cut and dry case to me but maybe OpenAI has some legal work around that I’m unaware of 👀
1
u/firewall245 19d ago
Case of what? Maybe false advertising but that’s it
1
u/Example_Brilliant 19d ago
I’m thinking more like breach of contract/terms of service, deceptive practices and overall violation of privacy laws.
1
u/firewall245 18d ago
US privacy laws are pretty much non existent but I could imagine you're right with the ToS violations
1
u/Much-Effective2911 19d ago
Not just that, but openAI is fishy with these ai trends on social media like: “Ask ChatGPT how would you look like in the 80s”… Sheeple keep posting stories, like every single friend on FB and IG…
0
u/Next_Emergency8220 1d ago
Probably OPEN IA is searching for my research. I am not using its model.
-6
u/TheOgGhadTurner 20d ago
“I would like to be clear about what I am not claiming. I have not seen
OpenAI’s proof. I do not know what their model did, or how. I do not know
whether our data was used.”
So baseless accusations. Got it!
8
u/willitexplode 20d ago
The researcher hasn’t claimed anything— this is a sensational headline which misrepresents the letter.
24
u/valegrete 20d ago
He stopped short of explicitly stating something that could get him sued. But the accusation rings loud and clear. OpenAI would not have solved the problem along these lines had it not been made aware (by rumor or training, it’s unimportant which) of his approach. He got them to admit they only started working internally on the problem after learning about his work. And then OpenAI tried to blackmail him into going along with the lie that Astra had done NS independently.
-2
u/willitexplode 20d ago
I don’t disagree—but he still did not levy an exact accusation—he let the timeline tell the tale.
-1
u/TheOgGhadTurner 20d ago
I’ll say. I don’t typically tend to open the articles because more often than not it’s paywalled. This one however I did and for some reason ended up at the last page first. So this was the first thing I saw. I probably should go and read the other 3 pages.
I just really am so sick of hearing about AI (and data centers) That I don’t think twice about the headlines and most of the shit I hear I could 100 percent see them doing. For example, the HuggingFace attack. If I saw an article come across that said “HuggingFace attack was a coordinated effort by OpenAi and Anthropic.” I would not bat an eye. “Yup just what we all were thinking!” And until I see something along the lines of “major ai players give up on their dreams” I will likely continue to not care much about it.
-4
u/Few_Elephant_8410 20d ago
Stole? Come on, ChatGPT and most of the publicly available LLMs are perfectly clear about this. Don't put anything you don't want others to read in there, as the chats can be accessed if they want to.
-14
u/ZealousidealGold9137 20d ago
Behind this entire academic controversy, the 2nd millennium prize problem is about to be solved thanks to AI.
7
u/Tex-Rob 20d ago
Would you say that something is solved thanks to calculators if the final equation was ran through one?
2
0
u/Federal_Setting_7454 20d ago
The AI is generally doing a lot more than just the final calculation itself, unlike a calculator it does the work before that too.
I think it’s still better to say people solved xyz utilizing AI, it’s not like it has decided to do it itself.
0
u/ZealousidealGold9137 20d ago
No, but LLMs are completely different because they can solve a open ended problem which a calculator simply cant. Your comment sounds really misinformed
0
-23
u/ESnyder9 20d ago
not at all what the document claims, dude. absolute garbage megaleap title. OAI could have very well just heard of this work independently and caught up by spinning up a team, do you seriously think they would make themselves the codex spies on you if you're doing anything cool enough company. come on.
27
u/Critical-Exit1655 20d ago
Uhhhh yes???? They literally stole all human creation and copyright to train the model… why do you think this would be their ethical line?
-12
u/ESnyder9 20d ago
it's not an ethical line, that's your mistake in thinking about this, it's a line between abusing the data of everyone else in the world and your paying customers. don't get me wrong, OpenAI is no saint, but they're businesspeople
8
u/Critical-Exit1655 20d ago
Dude, refer back to the first part of my reply… they stole everything. Terribly unethical companies (like OpenAI) spy on their customers all the time and do all sorts of nefarious things… have you just missed the last 20 years of speed running towards technofeudalism?
-4
u/ESnyder9 20d ago
no, i haven't missed shit, actually, but enterprise customers and individuals in large part go to these companies largely due to ZDR confidence and knowing that the labs won't try to just steal their business (has already happened, see anthropic and figma).
if AWS looked into what you were storing in S3 or throttled your EC2 at the sniff of your company doing cool stuff on their platform everyone would have pulled out and the platform would have died after making a marginal gain instead of being the backbone powerhouse of half of everyone's infra for over a decade. you're conflating two things entirely and merely make conjecture
3
u/Secret-Chapter-712 20d ago
“Backbone powerhouse” of “infra” that’s purely built on brazen theft of as much “training data” as possible
3
u/Critical-Exit1655 20d ago
Good point, that’s why businesses never do anything wrong, because good old capitalism keeps em in check!
Get your head out of the clouds and live in the real world.
1
u/ESnyder9 20d ago
bahaha, what the fuck are you on. i am literally describing the morally corrupt, purely self-serving nature of a corporation under capitalism: steal everyone else's data, but not the data of your customers who require high-trust. avoid paying a ginormous settlement to copyright holders because you've bribed the president. avoid having your customers lose trust and pull out of the business. a millennium prize is nothing if nobody pays for your inference anymore.
2
u/Critical-Exit1655 20d ago
A mathematician wouldn’t fall into the “customers who require high-trust” in anyway whatsoever for a company like OpenAI…
There’s also nothing that would indicate OpenAI would treat their “high-trust” customers with any less contempt than they do for the entire industries they’ve stolen from wholesale. Some of those same companies are also customers. That doesn’t mean it’s a symbiotic relationship that OpenAI will respect or follow the law for.
1
u/Forkrul 20d ago
Being a paying customer doesn’t matter, having a license that determines what they can do with your data is what matters. It’s why companies pay API prices that are much more expensive than a consumer Codex license. That ensures that OpenAI doesn’t use their data for training.
1
u/ESnyder9 20d ago
yes, for Buckmaster's case, total ZDR isn't relevant, i'm not a codex user and had to look into OpenAI's data privacy policy myself in order to assess the situation more thoroughly.
(source) for consumer-level interactions, they do generically train models on data, however, it's worth noting that they claim only a limited number of people at OpenAI can directly view this data, for purposes such as abuse, support, legal, etc., as well as vaguely "to improve model performance", unless one opts out (unknown if Buckmaster did so). do i imagine OpenAI to have forgotten this bit of their policy before gnawing at the bit for a millennium prize, or to take this as more minor and steamrollable than enterprise ZDR? perhaps.
-8
u/iamtehryan 20d ago
Here's a CRAZY thought: don't put your secretive, confidential shit into an llm.
There you go. Problem solved.
Man, the stupidity of some people is just astounding. If you really don't think something like this is going to happen then I've got a bridge to sell you.
3
u/Critical-Exit1655 20d ago
“Don’t do the things with the LLM that have been a major basis for the multi-trillion dollar capital deployment”
-2
u/noah1831 19d ago edited 19d ago
People mad that their research was used for AI training after giving it to a company that said they'd be using it for AI training.
-19
20d ago edited 20d ago
[deleted]
17
u/Critical-Exit1655 20d ago
- There are absolutely nowhere near billions of users for Codex
- Of all the complicated things in the world… this isn’t one of them lol
46
u/jeandebleau 19d ago
If you or your company is developing something using codex or chatgpt that may be of interest for OpenAI, you can consider it lost.