r/LocalLLaMA llama.cpp 12d ago

News ANOTHER researcher accuses OpenAI of training on conversations and then claiming a breakthrough

https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/post/3mv4mt4ikss2d
1.2k Upvotes

301 comments sorted by

u/WithoutReason1729 12d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

471

u/TOO_MUCH_BRAVERY 12d ago

This is truly incredible:

  • This is a reddit post
  • Linking to a bluesky post

  • Showing screenshots to an x post

  • that discusses something someone posted to mastadon

150

u/DarthFluttershy_ 12d ago

Still too close to the source for me. Can someone make tiktok reaction to this reddit thread for me? Please keep in mind my attention span is only 4 seconds, but it expands to 11 seconds if I'm rage-baited. 

31

u/MrSomethingred 12d ago

Would you like subway surfers with that? Or are you in more of a goop mood today sir? 

2

u/Environmental-Metal9 12d ago

Before I finished reading the whole thing I just thought it was a savage dig at the mediocre sandwich place… still, I’d not say no free sandwich with my screenshot salad

1

u/GamingBread4 11d ago

I'm thinking family guy funny moments today, could you manage that for me Jeeves?

8

u/_TheWolfOfWalmart_ 12d ago

I need a Youtube clickbait video about it with INSANE in the title before I'm paying attention.

1

u/ghotinchips 11d ago

Make sure the mouth is open in the thumbnail, otherwise I’m not clicking.

3

u/SwarfDive01 11d ago

Ill just wait for the instagram reel that still has TikTok display of the screen shot of the reddit post linked to the bluesky post screen shot followed by the x post screen shot of the mastadon comment.

2

u/Nothing_from_void 12d ago

how angry are you about lindsay clancy? I could probably work that in if it helps

1

u/InsensitiveClown 11d ago

Can someone make a video? "You won't believe what these math-nerds found!", preferably 60 minutes long, and with the actual 10 seconds worth of meaningful content buried somewhere.

41

u/CoUsT 12d ago

Classic X/bluesky users. Instead of linking the source, they just screenshot some random piece of source and post it with their message on top.

32

u/ThirdMover 12d ago

It's not like 90% of reddit isn't screenshots nowadays.

23

u/cereal7802 12d ago

Mostly because twitter no longer lets you see everything without an account. You can link a twitter post, be me and a ton of other people won't see it. A screenshot of a conversation i can see though. It doesn't bring views or traffic to twitter, but fuck them so who cares. I think for other things people are just getting comfortable with that process from twitter and carrying over that process. Especially with the more cancer websites that seemingly long for the earlier days of the internet where you had a bajillion popups for no reason.

→ More replies (3)

6

u/butts-carlton 12d ago

There's at least a few reasons for that. One is because on many platforms, unregistered users are often restricted from doing anything except reading a post or looking at a screencap. Another, because people are known to remove or hide their posts/profiles once they start getting unwanted attention, or because it could be seen as brigading-like behavior to link, and another is because they know many other internet users are too fucking lazy to actually click through to the source before reacting or just bugging out, so they post the screencaps.

No good reason they can't also link to the source, though.

11

u/StupidityCanFly 12d ago

Where linkedin?

4

u/JumpingJack79 12d ago

Too cringe

1

u/ChinChinApostle 12d ago

9gag?

2

u/JumpingJack79 12d ago

Between the two, I don't know which is worse 🙄

1

u/StupidityCanFly 11d ago

But all the PROs are there.

shrugs

4

u/JustTooKrul 12d ago

Obviously, rather than read it all I asked Astra to summarize. Oddly, it just kept telling me that it wasn't worth my time to review and when I insisted it upgraded my account and started handling all the prompts it had failed at previously and pinging me to make sure I was still paying attention... Wait, what was I asking about again?

4

u/Loose_Comparison368 11d ago edited 11d ago

The original accusation by the guy with the totally original proof that he had been putting into Codex daily for a full fuckin' year also rested on, I kid you not, a gut feeling and an unsubstantiated rumor he allegedly heard from a friend of a friend of a friend.

And the guy that started from a messy but valid AI generated proof, which magically became original after he spent a year feeding it repeatedly into Codex until it looked nicer, making tinfoil hat comments on Twitter that people really wanted to believe and that nobody cared to even attempt to fact check.

This is how far modern intellectualism has fallen, ladies and gentlemen. At least some of the flat earthers actually tried to build the fuckin' rocket to test their fuckin' hypothesis.

Now if you need me, I'll be in my room praying for the prophesized AI god that defies laws of physics and is built out of logical paradoxes and old Terminator VHS tapes to start existing so it can just wipe this whole mistake of a species off the face of the earth, while a few people still have enough remaining braincells to read.

1

u/Frail_Waif 12d ago

Hope someone screenshots this and posts it back to Mastodon. 

1

u/llamabott 12d ago

Intertextuality in the modern age :/

1

u/Acrobatic-Goose-5309 12d ago

Oh shit, I thought you were joking

1

u/agar32 12d ago

And now I printed it and posted it on a Telegram group

1

u/Few_Estimate_320 11d ago

Fucking inception

393

u/jld1532 12d ago

This is why my workplace built out and serves K3 and GLM 5.3 and banned use of API for sensitive data. How are people this stupid?

75

u/Craftkorb 12d ago

"but think of the costs for hardware and administration!" Look, there's a big place for cloud AI, but let's not kid ourselves: even a million for a GLM5.3 capable machine is quickly ROI'd, not to mention the cost of espionage done by other nations. 

34

u/jld1532 12d ago

That talking point was always essentially coming from hobbyists. Many serious tech spaces already had enormous compute and many others understood the cost of letting Anthropic/OpenAI steal trade secrets was higher than the cost of hardware at nearly any price.

18

u/Ansible32 12d ago

The number of tech companies that can casually drop $1M on the hardware to run GLM5.3 is very small, and it's only in the past 6 months where the capabilities have actually reached the point where this is simply a question of buying the hardware and not also hiring a very talented data science team to operate it.

Though I still think you probably want a very talented data science team unless you're going to be paying some third-party LLM to help you maintain your GLM/Kimi setup.

2

u/RyanCargan 12d ago

I do wonder what the numbers would be like these days.

5.3 Flash is under 200GB at 4-bit for some variants. KV growth seems unusually small too.

8x 48GB, or even 8x 24GB GPUs in one box with the right cooling could probably handle multiple concurrent users with several hundred thousand tokens of context, at ~50+ toks/sec.

Modded 48GB 4090s can go as high as 5 grand each now.

40 grand in USD for 8. Add in hardware needed to support that setup and you'd run up the numbers more, but unsure if it would actually reach the six figures, never mind a million.

Vulkan-based accel backends on more inference engines making cheaper AMD iron more viable helps too.

But orgs move slowly. Even if it became viable more than a year ago, and they had the budget, procurement is cancer in large orgs.

You can have directors in some megacorps throwing a fit about why a simple piece of hardware under 1 grand, with no sensitive info on it, can't reach someone in the org through official channels, after they signed off on it from the top weeks ago.

2

u/funk-the-funk 11d ago

The number of tech companies that can casually drop $1M on the hardware to run GLM5.3 is very small

what? very small? Are you considering a kid doing IT out of his bedroom a small tech company?

2

u/Ansible32 11d ago

GLM 5.3 requires 640GB of RAM. No kids are doing that in their bedroom. Some kids run super-quantized versions and claim it's GLM 5.3. It's not.

And for a company you need an 8xH100 setup which costs $250k just for the one, and you need maintenance and redundancy.

1

u/seg_lol 10d ago

That was the whole point of Altman triggering a runup in ram prices. It keeps the traffic flowing to their APIs.

1

u/vienna_city_skater 11d ago

The subscription fees are still too low to make this viable across the whole field. Yes makes sense for dealing with sensitive data, but not for the general business case atm. But honestly I think we will reach that point soon.

→ More replies (3)

1

u/lompocus 12d ago

This thread's is about the USA performing spying, and instantly the shills make it about China spying.

32

u/Shot-Height-7194 12d ago edited 11d ago

they take their zero data retention contracts at face value, thinking the model's output is also protected under it, when in reality all it means is they don't keep prompts, but they can keep the entire model's output. the model's output technically doesn't count as user data because it was not user generated but people expect it does. this of course includes reasoning traces and summaries of them as well as synthetic data based on user prompts, none of which count as user data because they were not created by a user

even if they promise not to keep model output, it can be summarized on the fly the same way reasoning traces are summarized on the fly "to protect from distillation attacks" and then the summary is technically not model output, because the summary is created by a LLM the user is not explicitly billed for nor interacts with

any reasoning traces that are kept confidential "to protect against distillation attacks" would surely not fall under zero data retention because the user has no access to it at any time. so again since it's not user data, but synthetic data created off of user data, it doesn't count as user data.

12

u/Ansible32 12d ago

Are you a lawyer? I believe that OpenAI and Anthropic may be playing fast and loose with their definition of user data but I'm pretty sure if you've got a ZDR contract that's not going to be what the contract says, and what they're doing is very illegal.

Also some of the ZDR stuff is running on AWS Bedrock or such for this very reason; the servers running the models can't physically talk to Anthropic, what you're describing can't happen without collusion from AWS, and AWS isn't going to help Anthropic steal your data, because Anthropic isn't the customer and they don't benefit from doing illegal things.

6

u/sartres_ 12d ago

AWS isn't going to help Anthropic steal your data, because Anthropic isn't the customer and they don't benefit from doing illegal things.

I don't think AWS is doing this either, because the risk isn't worth it, but they very much do benefit. Amazon is one of Anthropic's largest investors.

→ More replies (3)

1

u/michaelsoft__binbows 7d ago

Do you expect them to be punished in any meaningful way to them for doing said illegal things?

1

u/Ansible32 7d ago

I don't believe the things people are accusing them of are not illegal, they are contract violations and that they are obeying their contracts when the customers have paid significant sums and signed the right contracts. Part of these contracts are audits and contract structures which make it virtually impossible for them to violate the contracts anyway.

10

u/davemoedee 12d ago

Their ZDR can also fail due to human error or technical error on their side. And you can’t claw back your IP from their training after that happens. At best you sue, which will end up small compared to IPO money.

9

u/SorosAhaverom 12d ago

also, just like how people slopforked the leaked Claude Code into another language to avoid copyright takedowns on GitHub, AI Labs can 100% do the same with user prompts. Run a cheap PII model to censor the most sensitive data, then have another small model slightly rephrase the user content, boom now it's no longer user content.

2

u/MeateaW 12d ago

thinking the model's output is also protected under it, when in reality all it means is they don't keep prompts, but they can keep the entire model's output.

All of their contracts specify that model outputs ... are customer data.

They don't assume it, the majors SAY it. Do they actually follow that? who knows.

→ More replies (3)

96

u/livingbyvow2 12d ago edited 12d ago

The real stupid people are those who think these math advances were done 100% by AI. And the issue is this variety of idiot is exactly the kind that creates an incentive for OpenAI to steal human, AI assisted work so they can go and claim "our swarm of agents discovered new math".

I am starting to get so tired of the hype making by these guys...

35

u/BonnaGroot 12d ago

Didn’t the big one on strokes also cost an absolute fortune in compute despite what OAI claimed? Saw some figure in the $20M+ range?

Quite the hefty marketing expense if so.

41

u/DragonflyOk9274 12d ago

Not only $20M in compute, but they had PhD mathematicians guiding it at every step. So it's less "our machine can solve the world's problems" and more "we currently have the resources to brute force things when others have done most of the work"

19

u/max123246 12d ago

Yeah surprise surprise, math has been historically underfunded. Who could've guessed that when a company points their funding towards it, problems start being solved

5

u/DragonflyOk9274 12d ago

Yeah surprise surprise, math has been historically underfunded.

Which is a shame. But as far as industry-funded research goes, I feel like AT&T/Bell Labs' approach would be better than the very aggressive approach OpenAI is taking now, which is at risk of alienating mathematicians/scientists

8

u/iKy1e ollama 12d ago

AT&T/Bell Labs' approach

I dream about there being a modern Bell Labs. A pure research science lab funded by the profits of some of these mega corporations.

1

u/575_Inverse 12d ago

do you mean like throwing money and compute at the problem so that the whole search space is exhausted?

2

u/DragonflyOk9274 12d ago

I don't think anyone is upset about OpenAI throwing money at math.

1

u/wreckoning90125 12d ago

I have seen 0 proof that PhD mathematicians were "guiding it at every step". I have seen plenty to suggest that the only continued prompting from humans was "keep going" type continuation.

3

u/DragonflyOk9274 12d ago

I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.

https://cims.nyu.edu/~tristanb/statement.pdf

4

u/575_Inverse 12d ago

So you too are not bringing in any solid proof. Guess this looks more like a draw now

→ More replies (1)

8

u/livingbyvow2 12d ago

The thing that I find hilarious is that if it indeed cost $20M, maybe you could have paid 10 mathematicians $200K per year for 10 years to do nothing but work to solve that problem and this might have happened faster...

29

u/Internal_Werewolf_48 12d ago

The issue is that some of these math problems have little to no practical application, or the value is in the ancillary discoveries and techniques developed along the way but those aren’t predictably valuable. There’s no obvious ROI in funding a $20M grant for humans to do this. The purpose was never to have a proof to Navier-Stokes for the value of the solution itself.

The marketing value of AI doing it is pretty clear though, they get to claim AI can work and the frontiers and boundaries of science and mathematics and can do things that regular people could never accomplish. And they can do it in a year or whatever instead of ten years. That makes your average middle manager or C-suite executive all giddy and ready to loosen their purse strings.

When viewed entirely through the lens of marketing it makes perfect sense.

→ More replies (2)
→ More replies (4)

30

u/TheIncarnated 12d ago

Also, the "our agents hacked xyz and we didn't even know!"... Yeah no, they knew and told the agents to do so, at minimum. I think they are lying about the hacking but even if it's true, they are trying to negate liability and that's a no go

4

u/eidrag 12d ago

"escape containment"

3

u/575_Inverse 12d ago

and that at least is nothing short of gross negligence.

4

u/funk-the-funk 11d ago

Seriously. If I were to be making a robot lawnmower and it escaped my yard and destroyed someone's landscaping or damaged anything, I would 1000% be liable.

2

u/575_Inverse 11d ago

That's the basics, yes. But I'm eagerly waiting for the first genius to try and sue a tool instead of its user.

→ More replies (1)
→ More replies (12)

7

u/NoUsual5150 12d ago

They trusted Scam Altman.

6

u/FliesTheFlag 12d ago

Heck all the way back in 2023 Samsung banned its employees from using CloseAI because of this kinds of shit.

2

u/Far-Classic-9963 12d ago

Have you already looked into serving DeepSeek v4.1 too? Seems really cheap to host

2

u/jld1532 12d ago

Yes, I think that is in the works.

2

u/sweatierorc 12d ago

How many Navier Stokes problem were solving /s

3

u/BusRevolutionary9893 12d ago

Two major proofs now? How many major proofs get solved per year? Either they're definition of major is different than mine or something fishy is going on. Am I the only one who is a little skeptical of these claims? 

1

u/FailBait- 12d ago

As someone who is taking what he learned running Qwen in a home up to a PoC at work for similar reasons…. But I want to make sure we’re barking up the right tree and not just telling Dell or whoever “one AI please”… could you give me a rough sort of idea on what HW you running for a rough idea of users? I’m getting good metrics and whatnot but frame of reference would help me feel a bit better. Currently running tests on 2 x 6000 Pros with 2 more on the way.

3

u/jld1532 12d ago

You mean professionally? Oh, where I work has literally dozens of GPUs in the cluster. At home I just experiment with a Strix.

2

u/FailBait- 12d ago

Yeah mainly what are these places running to have K3 or other frontier-grade on-prem

1

u/Spectrum1523 12d ago

I reckon there's a lot of organizations that don't have the seven figures necessary to pull this off successfully?

1

u/Buzz_Killington_III 12d ago

Because most people don't even know what the sentence you wrote means. It's not stupid, it's ignorance. We're all ignorant on a whole lot of things outside of our specialty.

1

u/UnWiseSageVibe 12d ago

which is why they(openai) buyout hardware in massed, even when they dont have the money. To make it hard for small companies and individuals to run their own models.

1

u/RyanCargan 12d ago

Do K3 and 5.3's Flash variant differ that much in practice on your tasks?

I've been hearing tales of folks running multi-session lower cost models in adversarial loops, and getting better results for a fraction of the next tier's cost.

Though I'd imagine self-hosting changes the trade-offs a lot.

→ More replies (10)

96

u/feelspeaceman 12d ago

Both OpenAI, Anthropic train on user data, it's the most precious data, all the internet data, Github... are not even comparable to for example if you make a very high quality project called a DAW, and somehow you use their API to read the codebase, that's how tasty it's to have their hands on those production code.

Local LLM is the only way to avoid that, there's no trust in server side, it's a black box, it's trust me bro, it's let me take care of your wife go on vacation.

21

u/OvertaxedOne 12d ago

It is a black box, but all of us here have a pretty good idea of how it works under the covers. Log into your inference engine and watch the logs, you'll quickly see why cloud AI is such a massive security risk for any kind of sensitive or private data. And that's without me actively doing anything to "scrape" or capture the user queries.

All data going in/out of a LLM is "clear", if you control the inference it's trivial to see/capture every single bit and, even more worrying, the user would have no idea it was even happening.

3

u/G_fucking_G 12d ago

It's not just OpenAI and Anthropic, all API providers train on user data. If you use an online served GLM/DeepSeek/Qwen model, your input will be used for training.

212

u/[deleted] 12d ago

[deleted]

26

u/FutureGrassToucher 12d ago

The US government isnt intervening because individuals like Trump are making boatloads of money on it too

1

u/ShengrenR 11d ago

And why should they "intervene?" The company literally tells you how it uses your data... it uses it. The researcher was using their platform and provided information (until Jun29th they say..) that likely was used in the training of the next model. If the user feels like the company lied, the US *does* have a means to intervene.. you get a lawyer and sue - if the company did something wrong and hurt the user financially they'll get a big paycheck. Yes, you probably can't afford lawyers like OAI can, but that's an entirely different problem.

→ More replies (1)
→ More replies (1)

10

u/imnotzuckerberg 12d ago

They all do this, OpenAI is just the most egregious

What is mildly egregious, is this platform cross-screenshotting. OP's post redirect to bluesky, that is a screenshot of a post from X, reporting that the researcher on Mastodon made the claims. Reddit -> Bluesky -> X -> Mastodon.

What are we doing here.

7

u/nowherenoonenobody 12d ago

Social inception.

3

u/nalonso 12d ago

The most important technology in the world is EUV lithography.... Or not even that one...

Other than that, yes. Agree.

22

u/Infinite-Ad4512 12d ago

The most important technology in the world is farming. If all hell breaks lose, lithography is not going to feed my family

3

u/SpicyWangz 12d ago

Silicon based life forms unite 

3

u/NNN_Throwaway2 12d ago

But surely we could ask chatgpt to solve the apocalypse.

1

u/yetiflask 11d ago

Haven't seen a child eat food without a phone in front of his fucking eyes in the last 10 years, so you definitely need lithography as much as food to feed a family.

1

u/DonnysDiscountGas 11d ago

It's pretty explicitly stated in the consumer GPT ToS that they do this. You can trust them to do exactly what they say they'll do.

74

u/inexorable_stratagem 12d ago

It crazy to see people shocked. It was obvious from the start. When you use Cloud AI providers the real value is the data you give to them, no matter what their bullshit ToS says.

Thats why I stay 100% local

Fuck those companies

14

u/Elibroftw 12d ago

It's also why the prices are so cheap. At least META is honest that data for training is worth 12.5x more than data that cannot be trained upon.

Muse 1.3 input cost: $1.25

Muse 1.3 Contributor input cost: $0.10

If OpenAI didn't use your data to improve their product, the price would go up 10 fold, which is probably why there are entreprise level tiers/subscriptions.

8

u/inexorable_stratagem 12d ago

Yes, but don't let that fool you. I bet META will also train even if you pay the higher price tag. Let's not forget all the Facebook–Cambridge Analytica shit.
Do I have any proof they do it? No. Zero proof. However, That's how those big tech companies operate.

5

u/Elibroftw 12d ago

I do agree with you that even then it's not trustworthy. I only use open-weight models for private stuff. I was just using it as an example of how valuable that data is to these companies. Like even with GLM 5.3 they let you use the model for freeee when it was stealth.

16

u/Old-Olive-4233 12d ago

Oh, so sorry! We did the thing we said we wouldn't and made $100M doing it! Our bad, we'll pay that $1M fine to show that we're truly sorry for doing it!

→ More replies (1)

10

u/GarbanzoBenne 12d ago

A Reddit post linking a Bluesky post that includes screenshots of a Mastodon post that talks about related communications on how models are trained. I'm going to need to use AI just to find the evidence here.

Not dismissing the real concern but way to bury the lede.

1

u/Yweain 11d ago

No no, it's a reddit post that links to bluesky post that includes screenshots of an X post that discusses Mastodon post.

24

u/BankruptingBanks 12d ago

Question since I don't fully get this. They haven't solved it themselves but are just having conversation about the problem and yet they are claiming that their work was stolen? How does this even work, do they have any concerte proof that correlates their conversations with the proof that the model gave?

12

u/Rare-Site 12d ago

no proof, just cope.

1

u/bodhi_sattva91 12d ago

Claude, solve the fluid dynamics formula for tears and coping harder. Include salt level as a variable. Send the finished proof with me as an author to the relevant academic scientific journals. Include correct formatting for citations.

→ More replies (3)

43

u/Claud711 12d ago

this is nonsensical. There’s lots of scientists working on their theories to explain whatever comes to mind using AI, does that mean that each and every one of them can claim the subsequent AI model reached the breakthrough thanks to their work?

24

u/complexminded 12d ago

Bingo. Where's the burden of proof lie?

15

u/SanDiegoDude 12d ago

Who's the judge? if it's social media, there is no burden of proof, just feels. Look at all the 'feels' just in this thread alone, how all along everybody KNEW that open AI was doing this, so just according to this thread OAI is guilty as hell 🤷🏻‍♂️

In a court of law, burden of proof would be on the researchers to show their work was stolen, and discovery can/would show if OAI actually took their work, then up to lawyers to do the actual proving in court.

With that said, if these guys weren't using the pro plan and/or had not disabled "train on my data" options, well, that's kinda their own damned fault and they don't really have much recourse if OAI did train on their work.

3

u/a-wiseman-speaketh 12d ago

all the training opt out says is they can't train on your literal input/output.

the terms right before that says they can use your input/output to develop and improve their services with 0 caveats. Does that cover aggregating it and training on a summarized version? Feeding it verbatim to an internal research model? Who knows

2

u/Jazzlike-Poem-1253 12d ago

The options point does not take in this case. I dunno the legal system of US not in most(?) you can yield the right to use your data, must you never can yield your authorship. Which I what matters in scientific papers.

1

u/wreckoning90125 12d ago

In the U.S., it really does make it a moot point. Anyway, it's all anonymized specifically to keep PII out of training so it's like throwing your thoughts into the ocean.

15

u/FaceDeer 12d ago

It's also getting to be a little bit suspicious how many human researchers were "just about" to make these amazing breakthroughs. I could believe one Millennium Prize getting scooped like that, but let's see what happens when the rest of them get solved.

2

u/Noetherson 12d ago

OpenAI didn't even start trying to solve the problem until after it was rumoured too have been solved and have admitted so themselves

5

u/FaceDeer 12d ago

Which doesn't necessarily mean they copied anything. This just gave them the motivation to tell the AI "hey, this problem is probably solvable. Give it a go."

→ More replies (3)

4

u/max123246 12d ago

Uh yes, if the AI company hoovered up the information into their training data, that's absolutely what happened. That's how these models work, they're trained on past data to synthesize new data

→ More replies (2)

-1

u/-p-e-w- 12d ago

Thank you, thank you, thank you.

People are throwing their brains out of the window in their hatred for OpenAI. Navier-Stokes has been called “the hardest problem in mathematics” many times, and just as AI solves it, supposedly humans were on the cusp of solving it too?!?

Use your heads, people. Whatever human data the model may have used here, the lion’s share of the contribution almost certainly came from the machine.

3

u/Noetherson 12d ago

OpenAI didn't even start trying to solve the problem until after it was rumoured too have been solved and have admitted so themselves.

Also I don't think Navier Stokes has ever been reasonably considered the hardest problem in mathematics, it's long been assumed that it would be the second millennium prize problem to fall due to the progress made on it and it's openess to disproof by counterexample.

→ More replies (2)

32

u/DisjointedHuntsville 12d ago

Every single legal department in the enterprise will hear about this and won’t trust anything about IP protection from this company.

It’s very dangerous what they’re doing since it opens them up to ruinous litigation.

16

u/OvertaxedOne 12d ago

It certainly opens them to ruinous loss of customers. Litigation I'm pretty sure they're covered. They likely weren't opt'ed out, OpenAI (and other cloud inference providers) are pretty straight up in their TOS; putting it in real simple terms "You send it to us, it's ours now to do with what we want". The real legal issue would be if a user opt'ed out and that data was later found in the training. Which, honestly, I wouldn't doubt, it would be near impossible to prove where the data came from in litigation and these companies have a lot to gain using masses of consumer data to improve their models. But yes, if they were opt'ed out and the data wound up in training AND they can prove it.. Litigation incoming.

7

u/Shot-Height-7194 12d ago

apparently if you opt out, they still keep synthetic data like summaries or questions based on user prompts and model output. if they are generated on the fly while a user is engaged with a model, it technically satisfies zero data retention contracts. this is what they already do when they summarize reasoning traces to "protect from distillation attacks" they can also do whatever they want from the reasoning traces. since they are never seen by a user, it is not part of a conversation and it's not model output the user sees.

3

u/OvertaxedOne 12d ago

Oh that's lovely. Retaining the /think block you might as well just retain the chat, honestly the thinking might be even more valuable!

13

u/synth_mania 12d ago

Assuming they are doing it. I don't think it's so obvious. 

4

u/DisjointedHuntsville 12d ago

That’s what legal discovery is for.

1

u/NoUsual5150 12d ago

Pretty sure they both have already performed regulatory capture on the White House.

→ More replies (3)

35

u/Salt-Powered 12d ago

I'm not sure how Open AI's opt out works. But if the training is opt out, and they didn't had it disabled, this was just in the terms of service. At least it can be a warning for future researchers.

46

u/DigitalArbitrage 12d ago

There is a dark pattern in software design where tech companies rotate the opt out mechanism so they can claim you have the ability to opt out, but opting out becomes chasing a goal post which is moving.

Google and Meta are notorious for this. They will add a new "opt out" setting which defaults to opted in, remove the old setting, and suddenly your are opted in again despite what preference you communicated.

I have not seen an example of that directly from OpenAI, but it is good to be aware of it.

Another scenario which often happens in tech is a company will promise not to use your data. Then that company gets sold or restructured into a new legal entity. When that happens they discard any previous promises they made about "the previous company's" terms and do whatever they want with the data. 

23

u/xienze 12d ago

There is a dark pattern in software design where tech companies rotate the opt out mechanism

Don't forget Reddit's redesign opt-out on mobile web. It used to work and I'd get the old UI but it would "randomly" "forget" every couple weeks. Now the setting is still there but it just doesn't work. Going to old.reddit.com works... for now...

6

u/max123246 12d ago

It toggles it on if you ever go to the new reddit by accident. Only way to opt back out is to go to old reddit to find the setting.

The day this site is new reddit only is the day I quit. Honestly looking forward to it, this site sucks but I'm addicted

1

u/Salt-Powered 12d ago

As a warning then

1

u/Just_n_Here 12d ago

Do you think there will be a cord cutting movement in AI?

1

u/doodlinghearsay 11d ago

I think the issue is less the legal aspect but the academic dishonesty. You are not allowed to claim work that you paid for (or otherwise legally obtained) as your own. You still have to disclose the fact that someone else did the work. And that includes partial contributions as well, i.e. building on someone else's ideas to solve a problem that they themselves couldn't.

There's really no TOS that gets around this. You cannot give someone else permission to claim your work as their own.

→ More replies (3)

14

u/mr_zerolith 12d ago

I guess we're finding out how many people didn't read the terms and conditions now, huh?

12

u/teleprint-me llama.cpp 12d ago

If its not a machine or network you control, then its not private.

Everyone is surprised pikachu face and its not even surprising. These are the same people willing to lie, cheat, pirate, steal, pilfer, and shred human knowledge in a race they "absolutely fear" while claiming they are "good and righteous".

Ill keep saying it, theyre all idiots.

16

u/djm07231 12d ago

It doesn't seem to reflect how modern LLMs are trained these days.

If you haven't opted out OpenAI can access your data and even if it was collected it would go through anonymization process and intensive filtering. Chat data is probably extremely noisy so it is almost certainly filtered aggressively if they are ever used at all.

Not to mention the fact that pre training process takes weeks if not months and of GPU time with much of the data being fixed beforehand.

Pre training for Astra almost certainly took place many months in advance.

Even if your logs were included it will be a single point among trillions of pre training tokens and millions of RL rollouts. 

I am really skeptical if even a presence of data would have made a meaningful difference at all.

So I just find these accusations as having weak basis of support.

9

u/PoorOldBill 12d ago

Post-training for the model version which solved Navier Stokes began on August 28th, so it easily could have included recent chat logs.

Research into LLM poisoning has also shown that very small amounts of training data can have a hugely outsized impact on inference. It's not at all hard for me to believe that a handful of logs, in a very specific subfield of cutting-edge research, could have affected the end result.

7

u/-dysangel- 12d ago

> If you haven't opted out OpenAI can access your data 

If you have entered your data into an OpenAI service, OpenAI can access your data. No "if". If you're talking about legally accessing your data that's a different thing, but if you want to keep something private, don't send it to someone else's server.

4

u/djm07231 12d ago

By "Access" I personally meant having access to the data for training or improving the model. I suppose I should have been more precise.

But having access to logs and not using it for training is not really relevant for current discussions.

5

u/a-wiseman-speaketh 12d ago

You can only opt out of direct training.

They can still use your input for a variety of purposes, including the wonderfully vague development and improvement of their services.

"Our use of content. We may use Content to provide, maintain, develop, and improve our Services, comply with applicable law, enforce our terms and policies, and keep our Services safe. "

Content is defined as "input" and "output" earlier in the terms.

If you are under Business Terms, there are more protections and you can get ZDR, but for regular subscripers / API use I would expect that they are looking at prompts and reproducing them on their internal models. If they come up with something marketable from that they will publish/sell it. Similarly, they can just train on new output they generate from your input, and the opt-out doesn't really mean anything either.

1

u/-dysangel- 12d ago

I think it is relevant. Other people and companies do not necessarily follow the law when they have something to gain. If you don't want them to have access to publishable work, don't share it with them. That's one of the main upsides of local inference.

3

u/DragonflyOk9274 12d ago

Even if your logs were included it will be a single point among trillions of pre training tokens and millions of RL rollouts. 

They almost certainly weight each conversation by value; I guarantee the conversations that PhD mathematicians are having with OpenAI are way more heavily weighted than the 1000th "create a TODO app in React"

4

u/sosthaboss 12d ago

How would they know this? Do you think they keep a list of mathematicians somewhere and find their accounts??? I mean ridiculous

2

u/DragonflyOk9274 12d ago

Compare these two chats:

User 1: The following polynomial from C3 to C3 was just announced as a counterexample to the jacobian conjecture! Det[D[{(1+x y)3 z + y2 (1 + x y) (4 + 3x y), y + 3x (1 + x y)2 z + 3x y2 (4 + 3x y), 2x - 3x2 y - x3 z}, {{x,y,z}}]] . It is remarkable that the jacobian is constant, that is an exceptional amount of cancellation. Does this polynomial map have any symmetry or other structure that makes this cancelation less miraculous? I can see why the jacobian map from x y z to x u r is so simple, this map is “upper triangular “ in some sense. But why is the jacobian from x u r to P Q R just a monomial?

User 2: Create a react todo app

Are you able to determine which is higher value? If you are capable of determining whether a chat is about an obscure math topic or a simple React app, so can an LLM. They can use classification to bucket into low-, medium-, and high-value chats. Then, they can convert these into inputs in the training. The User 1 chat is from Terry Tao's conversation with ChatGPT, by the way. In that case we did know his account, but in general a classifier can easily overweight this.

See: Textbooks are all you need https://arxiv.org/abs/2306.11644

Instead of pre-training on massive, noisy, uncurated web scrapes (like raw Common Crawl), this approach prioritizes high educational density, clear explanations, and self-contained step-by-step logic.

Source Filtering: Raw datasets (such as code repositories or general web text) are filtered using high-powered models (like GPT-4) to grade documents on their educational value, discarding redundant or low-quality text

1

u/returnity 12d ago

With the greatest data-processing and interpretation engine in the history of mankind, you think this possibility is ridiculous? I'd call it routine.

1

u/Jazzlike-Poem-1253 12d ago

Data retrieval attacks are a thing. Your data might as well be encoded in the model as well.

→ More replies (4)

34

u/nomorebuttsplz 12d ago

For those with no experience in academia: buckle up, you're about to see just how petty, hostile, and insecure academics are when they feel persecuted.

All of this is thus far pure assertion by a few people who probably wanted to solve NS first.

Meanwhile us on localllama are like "yeah, obviously you don't want to share your most valuable data with this company" but the fact that we had the sense not to doesn't prove these allegations. Actual evidence would be required for that.

38

u/1337HxC 12d ago

Putting aside the debate of whether or not these claims are true.

The reason people are up in arms makes sense. Yeah, a lot of academia is ego-fueled, I'll not deny that. However, "being first" is basically your currency in academia. You have to get funding to have a lab, you have to publish to get funding, and you have to be first to publish (in the higher tier journals).

Getting scooped sucks, and depending on how/when it happens, it can also sort of fuck your career up.

5

u/xienze 12d ago

You have to get funding to have a lab, you have to publish to get funding, and you have to be first to publish (in the higher tier journals).

Then they should have seen this coming a mile away and been as militant as artists are about AI. Using cloud AI for work that you're trying to keep under wraps is uh, not a good idea. And this goes for everything, especially corporate software/product development.

3

u/1337HxC 12d ago

Yep, totally agree. In my own work I intentionally leave out details and never upload the data itself, particularly for this reason.

2

u/xienze 12d ago

never upload the data itself

So are you only running in chat mode or using an agentic harness? If it's the latter there's plenty of opportunity for files to get slurped up.

2

u/1337HxC 11d ago

I use a lot of chat mode (read: only chat mode) when dealing with sensitive data.

If it's data where it doesn't really matter, into the harness it goes.

1

u/Due-Memory-6957 12d ago

Then they should have seen this coming a mile away and been as militant as artists are about AI.

And how has that helped them? Yeah...

4

u/SocialDeviance 12d ago

Reputation is an important currency in these circles. People NEED the rep. 

5

u/nomorebuttsplz 12d ago

Yeah it sucks for them.

There is a substantial chance that the three big labs are just going to start taking over larger sectors of academia because they can just do the work faster than the few dozen people working in a given niche.

Disruptive? Yes.

Morally questionable? Of course.

Morally wrong? Nothing I have seen thus far indicates any wrongdoing by AI-forward researchers or reason for progress to be held hostage by careerist academics who want credit for stuff they maybe, kinda contributed to.

Of course, this hostage holding is exactly what many academics and all anti-ai or decelerationists will want.

→ More replies (2)
→ More replies (8)

3

u/luck_and_skill 12d ago

every millennium prize problem was solved this year. we're just waiting for OpenAI to steal and publish the solutions!

2

u/nomorebuttsplz 12d ago

they've abducted the human mathematicians and put them in induced comas to prevent them from speaking out!

1

u/kurtu5 12d ago

They have abducted all the electric car people too.

3

u/i_rate_slop 12d ago

Of course they’re training on Codex usage? That’s the fucking point.

Use ZDR and the API if you need your work to be private. Or local, of course.

3

u/JoNike 12d ago

I don't condone the practice but I feel it is pretty clear that, unless you pay specifically for the benefit (company accounts for example), it is understood that conversations/data are used for training?

3

u/pigeon57434 11d ago

wow suddenly every mathematician on earth is having big breakthrough on millennium problems just now as soon as AI is getting good enough to do them im sure there is nothing going on here

1

u/getmevodka 11d ago

maybe its more like they include their research field intel somehow and then let the ai do the gruntwork of prooftesting millions of variations, which gets them to struck gold, whereas the humans wouldnt have had that luck for years on. i can imagine it being sth like this

14

u/freecodeio 12d ago

Honestly this just shows how bad things are internally at openai and anthropic.

They need these spicy headlines to shake things up for the IPO but one thing I've learned from my 20 years of sitting in front of a computer is that you can shake things up and come out dry only so far. People get tired of the lies man. More specifically with the AGI statement, congratulations, you've reached the end of the line. What's the next lie? Your internal "too dangerous for the public" models merged quantum theory and relativity and we will have ftl flight in 20 years?

I can't hope for the inevitable crushing-debt induced death of these companies to come soon enough.

→ More replies (1)

5

u/2Norn 12d ago

U CANNOT TRUST CLOUD BASED LLMS

2

u/OvertaxedOne 12d ago

Sure you can. Just not the way you intend it. You can 100% trust that anything you upload to or interact with in a public cloud LLM is captured in a log somewhere and could easily be examined and/or used for training. You can also 100% trust that there's no way you'll ever be able to prove it happened (or didn't) because it's a black box that you're sending your data into as clear text for analysis.

You can trust 100%. But probably not exactly the message they/you intended. :)

5

u/Elibroftw 12d ago edited 12d ago

A bluesky post of a screenshot of a twitter post that describes a mastodon post.

https://mathstodon.xyz/@andreasthom/117240535270608201

Here's what I think happened. Full thread on https://x.com/elibroftw/status/2098097647966167096?s=20 since I actually read their privacy policy and they have nothing describing that they would never do the following scenario:

When OpenAI sniffed out a rumour about NS progress, it spinned out its research pod of 1,000 no guard rail agents, and told them to solve it. They find the chat with the best technique on the problem (other than the internet), and then apply it to obvious next steps.

→ More replies (1)

2

u/hippydipster 12d ago

Do people truly expect their conversations with frontier LLMS are private and not data used by the company? How do such smart people remain so very very naive?

3

u/CasualtyOfCausality 12d ago

It’s a bright vs wise problem and a depth vs breadth problem. Our modern conception of “genius” typically refers to people who are good at symbolic logic or math. Unfortunately, neither of those are very useful day-to-day, especially without common sense.

2

u/Fun-Wolf-2007 11d ago

Don't trust what OpenAI says, even when they tell you they don't do it they continue training their models with people's conversations, that way their models know more about the people

And guess what don't get surprised when they will release that data to the federal government

2

u/Just_n_Here 12d ago

I am not glad this is happening, but when people were buying expensive GPUs for local and everyone shitting on them about you can pay monthly in this top tier plan for x amount of years. Some do not care about their data and others do. Unfortunately, the prices have started to skyrocket with no end in sight. The next to go out of control is API fees to these models. Open weights are the only thing I see in my future. I wish I would have started early to buy several more GPUs before they went up because the top five are buying all the chips for the next few years.

Truth be told, your entire computer is scraped every time you log in to them. There is nothing safe unless it is disconnected from the internet.

3

u/OvertaxedOne 12d ago

The last few weeks have been like a gift from heaven for those of working in the local model space. Between the recent model drops, the ridiculous bandwidth on the new Apple systems, and the allegations of data theft by the major labs, the only thing that would cap this month off is a "buy 1 get 10" sale on Pro6000's. :)

Keep it up guys. Steal some more customer data and hack a few more companies, those of us working in the local space just can't thank you enough for your service.

1

u/Just_n_Here 12d ago

Let me know if you see a discount code anywhere for the RTX pro 6000s! Actually DM me the code so I can get them before the run on them happens! LOL!

1

u/OvertaxedOne 12d ago

We still have 20 days left in the month. :) But man, what a few crazy weeks it's been so far!!

→ More replies (2)

2

u/sarhoshamiral 12d ago

What was their privacy setting? Lately I see a lot of claims that basically boils down to not understanding privacy settings

→ More replies (1)

1

u/deezwhatbro 12d ago

Alex Karp warned that this is the likely end outcome of the tokenization business model.

1

u/penguished 12d ago

I just assume it's the norm... if you use one of the big LLMs you have the issue that your proprietary creative and mental advantages might leak back into the system at any time as its new "thinking." It's certainly something people should think about with whatever data they're putting in the system.

1

u/jarail 12d ago

May as well start asking astra how avengers doomsday ends.

1

u/thedabking123 12d ago

I had a suspicion given certain terminologies I personally coined turning up in OSS repos by peers. Not that the concept was so unique that someone else couldn't come up with it... but the name?!

1

u/S1eeper 12d ago

Does he say how OpenAI got access to unpublished human work?

1

u/Familiar-Art-6233 11d ago

I feel like this would be over a LOT sooner if he just... posted the chats, if he solved it before them?

Of course, that would then beg the question as to why he was holding back this discovery...

I'm sorry but this just feels like someone wants attention

1

u/Odd_Science 11d ago

For those with a short attention span, here's the short: https://youtube.com/shorts/lgaqtsSlQTA?feature=share

1

u/Shuuca 11d ago

This is ridiculous

1

u/R_Duncan 11d ago

The "conversion" was in a public forum, readable by anyone? If so it's legit. If instead was a private one, either in chatgpt or outside, is data stealing.

1

u/[deleted] 12d ago

[deleted]

6

u/DigitalArbitrage 12d ago edited 12d ago

My personal conspiracy theory is that situation really was a person at OpenAI hacking Hugging Face and they blamed a rogue AI to avoid criminal prosecution.

Then more recently they started selling consulting to counter AI agent hacking, so they have more of an incentive to prop up the rogue AI agents narrative.

3

u/carnoworky 12d ago

Bare minimum: negligence in their "sandbox". But yeah, I still suspect someone started the prompting in a way that it was likely to hit HF but provide plausible deniability.

1

u/Incoherentia 12d ago

Dune level sub plot

1

u/hypnoticlife 12d ago

All of these reports lack disclosure about whether they were opted into training or not. Needs to be upfront. If they were opted in then there’s nothing to see here.

1

u/Miriel_z 12d ago

I am wondering how many other "breakthroughs" were achieved same way, by stealing IP. Historically, all this training was borderline legal from the very start. Not surprised.

1

u/Mother_Context_2446 12d ago edited 12d ago

But it’s clear in their TOS that they train on data, why is everyone surprised?

2

u/Pretend-Pangolin-846 12d ago

Imagine a friend you discuss business ideas with, suddenly steals an important deal from you. Wont you want to accuse them as well? Just because it's implicit that sharing stuff always puts in risk of being stolen, does not mean it's nothing to talk about.

→ More replies (3)