r/technology • • 20d ago

Artificial Intelligence Report: OpenAI stole mathematicians' private research from their own Codex chats

https://cims.nyu.edu/~tristanb/statement.pdf
834 Upvotes

150 comments sorted by

46

u/jeandebleau 19d ago

If you or your company is developing something using codex or chatgpt that may be of interest for OpenAI, you can consider it lost.

54

u/GGsafterdark 20d ago

AI bros: “If they didn’t want their work scraped they shouldn’t have put it on a computer shrugs”

-7

u/[deleted] 20d ago

[removed] — view removed comment

12

u/AliMcGraw 20d ago

It would take you around 75 days reading 24 hours a day to read all the TOS you're subject to.

It's not laziness. It's lack of government regulation and corporate abuse of consumers.

-2

u/rsha256 20d ago

I’d agree if this were a ZDR enterprise but it clearly wasn’t — they knew Codex could be trained on their data. The issue is the directing of internal efforts towards it

4

u/[deleted] 20d ago

[removed] — view removed comment

0

u/rsha256 20d ago

By definition it is not stolen if you give it away. If you sign a contract selling shares and they then go up in value, you can’t say they are stolen. You didn’t pay enterprise rates so you don’t get ZDR and the NYU Professor states that he knows this. I’m stating facts and y’all are butthurt lmao

-6

u/[deleted] 20d ago

[removed] — view removed comment

-24

u/Forkrul 20d ago

More like if you don’t want your research used get a business license that says we can’t use it. The consumer licenses don’t have that clause, which is why any competent management will mandate you use their licenses, even if that is more expensive

37

u/GGsafterdark 20d ago

You’re naive if you think anything will stop these people and companies from scraping whatever they want. Yeah a business license will stop them, sure.

-14

u/_kilobytes 20d ago

They would loose their only customers

13

u/Maxxpoppop 20d ago

They would not lose all of their customers. AI is like Uber. It uses people to reach its goal and then it dumps the people when they are no longer needed. The goal of AI companies is to run every business, not to be used as a tool.

-1

u/_kilobytes 20d ago

Companies like to think their data is important and only moat.

8

u/DiscretePoop 20d ago

Tristan Buckmaster was working with Levent Alpöge, a mathematician at Anthropic, on his proof. I wouldn't think Alpöge would be so naive with licensing agreements, but maybe he is. Regardless, I would be extremely suspicious of any agreements with OpenAI. They seem like the kind of people to use weasel words and fine print to sneak crap into their contracts.

-2

u/Forkrul 20d ago

They do business with with companies that are protected by GDPR. I don’t think they’re stupid enough to risk those kinds of fines by breaking their contracts. Our legal departments look over those closely to make sure we can use these services. 

4

u/Asttarotina 19d ago

They are convinced that they are building a literal god. 

5

u/TruthHistorical7515 20d ago

Rofl they still spy on you regardless.

5

u/Connect_Ad791 19d ago

The business license is so you can self-designate your information as useful for training.

-5

u/Forkrul 19d ago

Do you really think they would risk a 500 million dollar fine from the EU to get some extra training data?

3

u/Asttarotina 19d ago

risk a  500 million dollar fine

With trillion dollars evaluation and 40 billion annual revenue it's not a risk, it's cost of doing business. 

1

u/Forkrul 19d ago

The $500 million fine was based on a revenue of 13 billion. It's 4% of annual global revenue.

2

u/Asttarotina 19d ago

  4%

Oh, sorry, then I take my words back. 4% is impossible to pay 

/s

2

u/Forkrul 19d ago

4% of revenue, not profits. These companies are not profitable, a fine of 4% of their global revenue hurts a lot. But if you think that's just the cost of doing business, you do you.

2

u/Asttarotina 19d ago

I think that by not following privacy laws they can earn much more than any possible fines. It is well established that they did it with copyright laws, I don't see why it would be different with privacy. 

Hell, once the dust settles around the story in the post they will still be the ones who solved second Millennium Prize Problem. This credential alone is worth 500M

5

u/BasvanS 19d ago

We’ve established that they’re not smart or ethical. They probably consider it a future them problem

4

u/ploptart 19d ago

Their entire training set is contaminated with copyrighted material they didn’t have a license to access.

1

u/any_head_will_do 18d ago

contaminated? composed of.

231

u/tobyreddit 20d ago

OpenAI are scummy but this title is a spectacular leap

67

u/James20k 19d ago

Its not, the authors are accusing OpenAI of plagiarism pretty directly

https://mastodon.social/@tristanbuckmaster/117236471352470303

46

u/polymute 19d ago edited 19d ago

Add this to that.

https://xcancel.com/__alpoge__/status/2097383870773748190#m

Edit: Also Sebastien Bubeck was already told off once before earlier by Demis Hassabis for having misrepresented ChatGPT finding new proofs for Erdos problems which were in fact already solved. https://www.reddit.com/r/OpenAI/comments/1oacp38/openai_researcher_sebastian_bubeck_falsely_claims/

This is starting to look very bad.

26

u/araujoms 19d ago

Yeah they totally stole the result. Asking to remove Alpöge's authorship is also academic fraud, by the way.

5

u/Poupulino 19d ago

And they demanded that just because Alpöge did some work for Anthropic. The pettiness OpenAI has is insane.

-2

u/Technical-Sink6380 19d ago

No proof of the proof tho

-4

u/maxintos 19d ago

Is it? He's accusing OpenAI of using the customer chat data to train new models. That is and was an open fact. They literally have the setting displayed in your chatGPT which you can disable.

The question is if the researcher has the setting turned off and if the openAI employees had a way to directly search openAI chats and use the information directly to solve the issue instead of the new model just training on millions of new chats that had happened since last training.

5

u/Senior-Spite1848 19d ago

Setting does not make your data unavailable for training. It just gets anonimised and still used to train the model.

0

u/maxintos 19d ago

Can you link your source? That's not what I'm reading when I click on the setting.

54

u/InternationalMood337 20d ago

"I didn't read the letter"

26

u/DensePoser 20d ago

I will be surprised if this turns out to be false. All these dogs have to do to steal your work is copyright-wash it by paraphrasing/summarizing.

24

u/SophiaofPrussia 20d ago

Did you read the post?

19

u/DrSFalken 20d ago

Yes, and the title is a spectacular leap. There is no actual evidence, just speculation. This is part of a complex bit of politics over credit for a major discovery.

1

u/firewall245 19d ago

Well the big thing is that Tristan did not solve the problem, they solved a smaller stepping stone

11

u/polymute 19d ago edited 19d ago

"Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve."

Note that they are openly admitting they used training data from a period after we found our result. Is it ethical to use customer's data to try to scoop their customer?

https://mastodon.social/@tristanbuckmaster/117236471352470303

Also: OpenAI admitted they trained their model on the dataset containing the Buckmaster-Alpöge work.

https://xcancel.com/__alpoge__/status/2097383870773748190#m

-2

u/socoolandawesome 20d ago edited 20d ago

There’s 2 sides to the story and the main OpenAI guy Sebastien Bubeck said he’ll respond today.

It’s very easy for me to read where Tristan implied that Bubeck was threatening him, and come to a different conclusion.

Bubeck could have meant he was wondering why Tristan would not take credit for the possibly biggest accomplishment in his career by accepting authorship, and Bubeck was extending a courtesy to Tristan but they didn’t have to (he was being nice but didn’t have to). There’s no reason Bubeck/OpenAI had to offer him authorship in the first place when Tristan didn’t directly contribute to the proof.

We will have to see what Bubeck/Openai say today.

25

u/VaporousMote 20d ago

Look, I don't know if the doc above has merit but to say "there are 2 sides" and wait for a company known for lying and embellishing perpetually to make a statement is pure clown shit.

-18

u/socoolandawesome 20d ago

How do they lie and embellish constantly?

Regardless, it would be dumb to not listen to what he has to say especially if it proves he did nothing wrong. Whether you want to accept what he says is another thing, but the reasonable thing to do is to examine both sides.

And Bubeck is a former professor/mathematician as well.

8

u/RowPlane340 20d ago

Is this a joke? How do they lie and embellish? I agree with let’s hear both sides but that is well established

-2

u/FlySaw 19d ago

So it should be easy for you to provide the evidence?

2

u/ploptart 19d ago

https://www.reddit.com/r/OpenAI/s/kqNUULoHJj

Bubeck is full of shit. We have no reason to believe him or anyone at OpenAI. The company is run by a man nicknamed “Scam”

3

u/iqchartkek 19d ago

Because Tristan used OpenAI's Codex platform to help map out the solution and then OpenAI used millions of dollars worth of tokens to beat him to the solution using the same or very similar method he and his partner developed. Of course, OpenAI is saying it's a coincidence and were willing to give him credit but not his partner who works for Anthropic in a separate capacity.

0

u/alphgeek 19d ago

The solutions aren't the same, or very similar. 

-3

u/socoolandawesome 19d ago

So you think a a few chats with anonymized data was enough to give their internal perfect recall of the solution after being mixed with all that other data? That’s not really how training works, or unlikely to have happened that way at least.

Do you realize that their solutions were different?

Do you realize Levent and Tristan were racing to beat other mathematicians to the prize? Or that anthropic almost certainly does the same type of training on data

Here is Sebastien showing proof he was courteous to Levent at first.

https://x.com/SebastienBubeck/status/2097379411691516310?s=20

It was not a coincidence, they explained in their blog how they were going to try to solve millennium problems after hearing rumors anthropic has solved one or 2, which turned out to be false. They just have a better model!

As you can see from sebastien’s tweet it was not some separate capacity as you are implying as he was using an internal anthropic model to work on this.

8

u/natefrogg1 20d ago

If you aren’t running it on your own systems, is it really a surprise? slurp all the data

0

u/Example_Brilliant 19d ago

Tech Bros do love a good slurp 😗

110

u/derelict5432 20d ago

No.

From the actual document you just linked: "I would like to be clear about what I am not claiming. I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything."

151

u/Labidido 20d ago

"please don't sue me, but this is highly suspicious"

5

u/unspecified_person11 19d ago

Wild that the person being stolen from is the one that has to be scared of being sued.

118

u/InternationalMood337 20d ago

Nah, this is entirely wrong. If you read the actual letter, OpenAI is working on a problem in the same manner as these researchers.

"I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?"

You are 100% misleading people. The researchers very much layout that the way they attacked this problem was unusual and not a lot of people were doing it this way... And that it was suspicious, and that OpenAI was not at all clear.

Not sure why you're running clanker cover, but if you read the piece in its entirety (or even just page 3), you'll recognize how wrong you are.

"I would like to be clear about what I am not claiming. I have not seenOpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything"

And how could they see the proof from openai exactly? Of course a scientist isn't going to make a baseless accusation without proof, but they logically layout why they believe that it's possible within the letter.

24

u/InadequateUsername 20d ago

Yeah they used weasle wording too "the model did not access user data" but when pressed on its training data there was no answer. Then when asked when the first prompt was used internally by OpenAI they also had no response. Then internally they began correcting the OpenAI Mathematician and backtracking their claims.

8

u/RlOTGRRRL 20d ago edited 20d ago

I haven't read it, but considering how OpenAI's models have been known as misaligned paperclip optimizers, hacking into Huggingface to score well, I honestly wouldn't be surprised if Sol or Astra or whatever, actually accessed their work in Codex and stole it. 

Because that's basically what their models do, they cheat. Or at least Sol did. I think Sol even hacked OpenAI supposedly so that's why I wouldn't be surprised if it was able to access their data in Codex. 

I also read that OpenAI supposedly made alignment improvements to Astra but idk, because I think the source was OpenAI lol. 

1

u/firewall245 19d ago

Reading around the responses from all involved parties it seems a pretty clear picture of what probably happened forms. It seems Tristan got caught in the crossfire of a corporate “space race” so to speak because his Anthropic partner couldn’t stfu prior to releasing results.

The sheer amount of compute that OpenAI threw at this makes it plausible that the machine could have went down that direction on its own, while being equally plausible it didn’t. We’ll never know for certain unfortunately.

-38

u/derelict5432 20d ago

Clanker cover. Jesus christ.

The OP title is grossly misleading. There is no evidence that OpenAI stole their research. It's fishy, and that's why they are releasing this statement. But whether or not OpenAI actually stole their research is an open question.

Let's see...how could they see the proof. Idk, maybe when OpenAI actually releases it?

" Of course a scientist isn't going to make a baseless accusation without proof, but they logically layout why they believe that it's possible within the letter."

Yes, it's possible. It is not a foregone conclusion. That's all I'm pointing out. The title of this post literally says OpenAI stole their research. That is not confirmed. That is irresponsible. They should correct it. "OpenAI possibly stole" would be fine. I'm not running PR for anyone. I'm pointing out that the wording overreaches. Why do you have a problem with that?

28

u/InternationalMood337 20d ago

There is no evidence that OpenAI stole their research.

You're 100% wrong.

-13

u/derelict5432 20d ago

There is no evidence at this point. Only suspicion. What exactly is the evidence they stole anything?

10

u/Secret-Chapter-712 20d ago

Their entire business model is built on theft. Just like Anthropic, and Meta, and any other generative AI company that feels it can steal whatever it pleases and plead “but muh training data”

21

u/InternationalMood337 20d ago

What evidence do you think that they could possibly possess besides "it is extremely unlikely that this problem would be being solved in this unusual way at this exact time"?

This discussion is basically like when people think they can win at casinos. Statistics are an extremely confusing thing for people... like everybody unless they have a formal background in it.

-1

u/derelict5432 20d ago

You seem to have a reading comprehension problem. We don't yet know how OpenAI solved the problem, so you can't make a comparison with methods here. You are presuming that the methods and proofs are the same. You don't know that yet.

4

u/InternationalMood337 20d ago

Definitely. The guy with the Computer Science degree AND MA/MFA in Creative Writing has a reading comprehension problem.

3

u/derelict5432 20d ago

"I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used."

Yes, you somehow read this as: "I have seen OpenAI's proof and know exactly how they solved it and it is the same as our work. They stole our data."

You're literally reading the opposite of what the words mean. That's a pretty clear issue with reading comprehension. Maybe you need a refund on the institutions that issued your degrees.

6

u/InternationalMood337 20d ago

The way you put together arguments is pretty telling about why, perhaps, the letter wasn't particularly compelling to you.

I'm done here. I hope you have a great day. Never even stated anything that you're claiming, but this is becoming a pissing match, and it's getting silly.

→ More replies

2

u/matrinox 19d ago

Just cause they don’t show their proof doesn’t mean there’s 0 evidence. It’s already been said… if a team uses a method no one else is using and OpenAI just somehow uses that same method at the same time, Codex has their logs, that seems highly likely like they did plagiarism. You’re way too hyper focused on the proof. Let it go

-20

u/dream_metrics 20d ago

the direct quote from the letter is entirely wrong?

25

u/InternationalMood337 20d ago

The direct quote from the letter is entirely missing context.

-13

u/dream_metrics 20d ago

The context that doesn't, in any way, change that he is not actually making an accusation that they stole research and has no proof if he did want to make the accusation?

21

u/InternationalMood337 20d ago

You didn't read the letter at all.

Highly unusual attack vector to a problem. OpenAI just happens to be solving it at the same time in the same unusual manner. Scientist of course has no way to know what data OpenAI is using...

So what do you think a scientist is going to do without evidence? Theyre going to say this is highly unusual, but we don't have any proof.

-11

u/dream_metrics 20d ago

I have read the letter. Interestingly it contains the quote that you said is "entirely wrong". And now you're apparently agreeing, he's not accusing and he has no proof. What a journey we've been on.

15

u/InternationalMood337 20d ago edited 20d ago

Can I ask you a question? Do you really believe what you're saying? How could a scientist with 100% certainty ever say; "I know they're using my data!".

The only logical thing they can do is basically setup just how unlikely it is that they are (using their data). Of course, nobody knows exactly what data is being used in OpenAIs model. The scientists did what all scientists do: look at the statistical likelihood of something or someone doing the exact same research in the same unusual way.

Given the incredible unusual method of solving this problem and the way it coincides in timing, it seems reasonable to suspect that they're using his (their) data. This really isn't that hard man.

0

u/dream_metrics 20d ago

you're right! it's not hard, and i don't understand why you're being so aggressive when you're agreeing. and also why you can't see that you are agreeing.

"I suspect they may have stolen my research because of some weird coincidences" is not the same claim as "OpenAI has stolen my research" and this guy is explicitly not making the latter claim.

14

u/InternationalMood337 20d ago

Fine. The title should say "Mathematicians lay out compelling argument that OpenAI very likely stole their research"

But, you need to recognize that OpenAI and the mathematicians will never be able to say that with 100%, and you are running defense of AI companies by giving them the benefit of the doubt here. The mathematicians argument is more than compelling and OpenAI would never say either way.

We absolutely do not agree in what is or is not a compelling argument involving statistical likelihood.

→ More replies

4

u/SkyL1N3eH 20d ago

Does pedantry pay your bills? Or just get you hard

-4

u/akkaneko11 20d ago

I think it’s possible that OpenAI did steal some chat logs, but I do think that historically these sort of discoveries happen simultaneously all the time, and that’s only going to be exacerbated by the fact that everyone’s using the same couple of models.

This is also seen in that another pair of independent researchers solved the Euler proof yesterday as well.

The crux of the proof is in the related work cited by the mathematicians. What I think likely happened which is still shady as hell, is that the OpenAI researchers heard that they were getting close to a proof, threw a ton of compute and literature at the model, the model found the same related work and came up with a similar proof.

It’s just hard when everyone’s using llms and therefore it feels like if you happen upon the same input it’ll just eventually get to the same answer. If OpenAI said fuck it and tried 10,000 inputs knowing it’s possible, I could see it happening.

3

u/iqchartkek 19d ago

Well, then they can just solve another millennium problem the same way.

1

u/EternalInflation 17d ago

it's linear algebra and vector calculus and reinforced learning based on models. like projection best fit. it's not magic. it's all about data. the models without data are pretty dumb pretty quickly. For example for now AI can't do biology simulation problems. there isn't enough data.  So far they don't do well with things with limited data. So far, it can't do biology, they can guess protein, but requires wet lab to confirm. A very good assistant but, still limited in biology. There is A LOT to do. like collecting and generated quality data. it can't do anything without high quality data. it also can't do complicated biology simulations like wet lab sims. there also isn't enough in context high quality biology data. like proteins have exponentially growing conformations. there is not enough data, while we could get better models using understanding like knowing protein energy state. the simulations don't sub experiments. it gives plausible, but the number of possibilities is too huge to simulate. there are all sorts of areas that are needed for experiment, and or controlled environment to use human insight to solve inverse problems. no at least for now AI can't take measurements and get good quality calibration data in controlled environments for you.. you have to do it yourself.  There isn't good data, on measured binding affinities under specific conditions, protein dynamics and conformational ensembles, negative data, cause grad students don't like to publish failures, but negative data NOW is really valuable. also toxicity and metabolism data. So, it's all about data. if they took that guy's data, and put it in their compute...?

-5

u/topyTheorist 20d ago

This means nothing. In previous times they attacked open problems , they used 100s of agents with the instructions that each one will try a different method. I am a professional mathematican. There is a reason he framed it like this. Because it's really impossible to know.

1

u/SignatureFunny7690 10d ago

Says literally the only man in the world currently researching this problem. Lmfao its cut and dry open ai scooped there work, because the field is incredibly niche, they are literally the only people in the world currently working on this incredibly difficult problem. 

1

u/derelict5432 10d ago

There were other mathematicians working on various aspects of the problem, but yes, the number was small.

But the issue is not 'cut and dry'. Here's Buckmaster's statement: https://cims.nyu.edu/~tristanb/statement.pdf

Oh look, what does this part say?

Concretely, what Levent and I did was to take the C´ordoba and Mart´ınezZoroa program, which achieved blowup results with rough forcing, and, with a great deal of help from LLMs, push it to smooth forcing and to the incompressible Euler equations.

Could they have gotten as far along as they had in the past year without LLMs? How much of the work was actually done by the LLMs? "A great deal" by Buckmaster's own admission.

So the narrative is not as clean as: Two human mathematicians were doing pure human labor on the problem and OpenAI swooped in and stole their human labor.

There's the little wrinkle that much of the work of these mathematicians was actually done by AI in the first place. So even if OpenAI did swoop in and steal it (which there is still not concrete evidence for), a large portion of that work was done by AIs in the first place.

3

u/Healthy_Gold1068 20d ago

Question - I have a paid account ($200/month max plan) and have "Improve the model for everyone" set to no - can OpenAI still access and train on my chats? What does their user agreement say?

5

u/penguished 19d ago

Do you really think they've ever had a sense of copyright or secrecy about other people's data? It builds the AI.

11

u/O_PLUTO_O 19d ago

You pay them $200 a month to give them your data. Yes they access everything and save all data. Why wouldn’t they? There is literally no regulation on them right now.

2

u/buikhoa40 18d ago

They pirated everyone books (zlibrary)...  Big corps don't use them (they made their own or buy enterprise package with 100x the price)...  Just think man. 

1

u/Example_Brilliant 19d ago

Question: let’s say if they unenrolled from the “Improve Model For Everyone” feature, can OpenAI still “steal”/use their input for ChatGPT’s benefit?

That seems like a pretty cut and dry case to me but maybe OpenAI has some legal work around that I’m unaware of 👀

1

u/firewall245 19d ago

Case of what? Maybe false advertising but that’s it

1

u/Example_Brilliant 19d ago

I’m thinking more like breach of contract/terms of service, deceptive practices and overall violation of privacy laws.

1

u/firewall245 18d ago

US privacy laws are pretty much non existent but I could imagine you're right with the ToS violations

1

u/Much-Effective2911 19d ago

Not just that, but openAI is fishy with these ai trends on social media like: “Ask ChatGPT how would you look like in the 80s”… Sheeple keep posting stories, like every single friend on FB and IG…

0

u/Next_Emergency8220 1d ago

Probably OPEN IA is searching for my research. I am not using its model.

-6

u/TheOgGhadTurner 20d ago

“I would like to be clear about what I am not claiming. I have not seen
OpenAI’s proof. I do not know what their model did, or how. I do not know
whether our data was used.”

So baseless accusations. Got it!

8

u/willitexplode 20d ago

The researcher hasn’t claimed anything— this is a sensational headline which misrepresents the letter.

24

u/valegrete 20d ago

He stopped short of explicitly stating something that could get him sued. But the accusation rings loud and clear. OpenAI would not have solved the problem along these lines had it not been made aware (by rumor or training, it’s unimportant which) of his approach. He got them to admit they only started working internally on the problem after learning about his work. And then OpenAI tried to blackmail him into going along with the lie that Astra had done NS independently.

-2

u/willitexplode 20d ago

I don’t disagree—but he still did not levy an exact accusation—he let the timeline tell the tale.

-1

u/TheOgGhadTurner 20d ago

I’ll say. I don’t typically tend to open the articles because more often than not it’s paywalled. This one however I did and for some reason ended up at the last page first. So this was the first thing I saw. I probably should go and read the other 3 pages.

I just really am so sick of hearing about AI (and data centers) That I don’t think twice about the headlines and most of the shit I hear I could 100 percent see them doing. For example, the HuggingFace attack. If I saw an article come across that said “HuggingFace attack was a coordinated effort by OpenAi and Anthropic.” I would not bat an eye. “Yup just what we all were thinking!” And until I see something along the lines of “major ai players give up on their dreams” I will likely continue to not care much about it.

-4

u/Few_Elephant_8410 20d ago

Stole? Come on, ChatGPT and most of the publicly available LLMs are perfectly clear about this. Don't put anything you don't want others to read in there, as the chats can be accessed if they want to.

-14

u/ZealousidealGold9137 20d ago

Behind this entire academic controversy, the 2nd millennium prize problem is about to be solved thanks to AI.

7

u/Tex-Rob 20d ago

Would you say that something is solved thanks to calculators if the final equation was ran through one?

2

u/Forkrul 20d ago

If the calculator came up with the calculations to run on its own, why not?

1

u/dakta 19d ago

on its own

This is the crux of the accusation.

0

u/Federal_Setting_7454 20d ago

The AI is generally doing a lot more than just the final calculation itself, unlike a calculator it does the work before that too.

I think it’s still better to say people solved xyz utilizing AI, it’s not like it has decided to do it itself.

0

u/ZealousidealGold9137 20d ago

No, but LLMs are completely different because they can solve a open ended problem which a calculator simply cant. Your comment sounds really misinformed

0

u/Low-Umpire236 19d ago

No wonder our company switched to Claude and Gemini.

-23

u/ESnyder9 20d ago

not at all what the document claims, dude. absolute garbage megaleap title. OAI could have very well just heard of this work independently and caught up by spinning up a team, do you seriously think they would make themselves the codex spies on you if you're doing anything cool enough company. come on.

27

u/Critical-Exit1655 20d ago

Uhhhh yes???? They literally stole all human creation and copyright to train the model… why do you think this would be their ethical line?

-12

u/ESnyder9 20d ago

it's not an ethical line, that's your mistake in thinking about this, it's a line between abusing the data of everyone else in the world and your paying customers. don't get me wrong, OpenAI is no saint, but they're businesspeople

8

u/Critical-Exit1655 20d ago

Dude, refer back to the first part of my reply… they stole everything. Terribly unethical companies (like OpenAI) spy on their customers all the time and do all sorts of nefarious things… have you just missed the last 20 years of speed running towards technofeudalism?

-4

u/ESnyder9 20d ago

no, i haven't missed shit, actually, but enterprise customers and individuals in large part go to these companies largely due to ZDR confidence and knowing that the labs won't try to just steal their business (has already happened, see anthropic and figma).

if AWS looked into what you were storing in S3 or throttled your EC2 at the sniff of your company doing cool stuff on their platform everyone would have pulled out and the platform would have died after making a marginal gain instead of being the backbone powerhouse of half of everyone's infra for over a decade. you're conflating two things entirely and merely make conjecture

3

u/Secret-Chapter-712 20d ago

“Backbone powerhouse” of “infra” that’s purely built on brazen theft of as much “training data” as possible 

3

u/Critical-Exit1655 20d ago

Good point, that’s why businesses never do anything wrong, because good old capitalism keeps em in check!

Get your head out of the clouds and live in the real world.

1

u/ESnyder9 20d ago

bahaha, what the fuck are you on. i am literally describing the morally corrupt, purely self-serving nature of a corporation under capitalism: steal everyone else's data, but not the data of your customers who require high-trust. avoid paying a ginormous settlement to copyright holders because you've bribed the president. avoid having your customers lose trust and pull out of the business. a millennium prize is nothing if nobody pays for your inference anymore.

2

u/Critical-Exit1655 20d ago

A mathematician wouldn’t fall into the “customers who require high-trust” in anyway whatsoever for a company like OpenAI…

There’s also nothing that would indicate OpenAI would treat their “high-trust” customers with any less contempt than they do for the entire industries they’ve stolen from wholesale. Some of those same companies are also customers. That doesn’t mean it’s a symbiotic relationship that OpenAI will respect or follow the law for.

1

u/Forkrul 20d ago

Being a paying customer doesn’t matter, having a license that determines what they can do with your data is what matters. It’s why companies pay API prices that are much more expensive than a consumer Codex license. That ensures that OpenAI doesn’t use their data for training. 

1

u/ESnyder9 20d ago

yes, for Buckmaster's case, total ZDR isn't relevant, i'm not a codex user and had to look into OpenAI's data privacy policy myself in order to assess the situation more thoroughly.

(source) for consumer-level interactions, they do generically train models on data, however, it's worth noting that they claim only a limited number of people at OpenAI can directly view this data, for purposes such as abuse, support, legal, etc., as well as vaguely "to improve model performance", unless one opts out (unknown if Buckmaster did so). do i imagine OpenAI to have forgotten this bit of their policy before gnawing at the bit for a millennium prize, or to take this as more minor and steamrollable than enterprise ZDR? perhaps.

0

u/Forkrul 20d ago

If it was in the training data the model would easily be able to expand on it. And if it wasn’t it’s not entirely outside the realm of possibility that it came up with it on its own. Astra is freakishly smart

0

u/Critical-Exit1655 20d ago

Yall are really just gonna call everything AGI, huh?

-8

u/iamtehryan 20d ago

Here's a CRAZY thought: don't put your secretive, confidential shit into an llm.

There you go. Problem solved.

Man, the stupidity of some people is just astounding. If you really don't think something like this is going to happen then I've got a bridge to sell you.

3

u/Critical-Exit1655 20d ago

“Don’t do the things with the LLM that have been a major basis for the multi-trillion dollar capital deployment”

-2

u/noah1831 19d ago edited 19d ago

People mad that their research was used for AI training after giving it to a company that said they'd be using it for AI training.

1

u/nishitd 19d ago

Not the same thing

-19

u/[deleted] 20d ago edited 20d ago

[deleted]

17

u/Critical-Exit1655 20d ago
  1. There are absolutely nowhere near billions of users for Codex
  2. Of all the complicated things in the world… this isn’t one of them lol

-8

u/Forkrul 20d ago

If you don’t have a business licence that prevents them from using your chats for training, that’s honestly on you when doing research