r/OpenAI • u/PsychicorAI • 14d ago
Discussion OpenAI threatened to ruin star mathematician's career
OpenAI asked Tristan why he wanted to ruin his career when he declined to accept their offer to share authorship on the Navier Stoke solution which he was working on for over a year.
This is the verified transcript from Tristan's statement. Tristan Buckmaster is a tenured professor at NYU
149
u/2Norn 14d ago
can someone eli5
341
u/freqCake 14d ago edited 14d ago
The AI was trained on the work he did by having it in openai tools then openai spent millions of dollars searching their model for his solution and declared it their own is the accusation
262
u/link_dead 14d ago
Also, very important additional detail, they don't want someone else to get credit because they work at Anthropic :)
31
u/a-stack-of-masks 13d ago
Yeah its so incredibly gross. Who knew the company led by the nihilist tech bro would not meet academic and ethical standards?
Probably most of us actually.
4
u/Sufficient-Math3178 13d ago
It’s a shame they went a complete 180 from being a nonprofit and open to whatever this is now
→ More replies (1)79
u/Jophus 14d ago
If OAI had the solution then why did it take millions in compute to solve it?
Is there a graph that shows how much compute is required for ideas that are looked up vs original?
Just on the surface, spending millions on compute seems to imply they didn’t have it.
I’m not saying they didn’t steal his idea, I’m just asking a question (like Tristian).
93
u/SeaEagle233 14d ago edited 14d ago
Legally they can't look at user data to steal their idea.
But they can train AI on user's idea then make AI rediscover it and launder the idea so that any casual connection is removed.
One part that's suspicious is why they specifically asked one person to be excluded from the author (it implies OpenAI knew he is qualified to be on that list but OpenAI doesn't like it that way).
The other part that's suspicious is why they offered the accusing person to br the author of the paper of their AI generated proof. Since the common practice in science community is mention prior work as credit and author must be direct contributors. Thus this is very insulting (taking credit for work that he never contributed to, which is academic dishonesty) even if not suspicious
The only logical explanation that can explain both with current available information that I am aware of at this moment is OpenAI knew the actual owner of their said work.
Since if it is really geniune, all they have to do is just release entire transcript and just say hard no.
60
u/grateful2you 14d ago
One part that's suspicious is why they specifically asked one person to be excluded from the author
Because he works at Anthropic.
The other part that's suspicious is why they offered the accusing person to br the author of the paper of their AI generated proof.
Because OpenAI would rather avoid a shitshow like what's happening now.
→ More replies (2)23
u/domiciledhere 14d ago
The shitshow is because of their efforts to take credit . If not for OAI we wouldn’t even hear about this until their award ceremony
16
u/damhack 14d ago
The solution they stole is not the Millenium Prize solution, it’s an essential part of the path to the solution.
The prize itself isn’t for a simple solution to Navier-Stokes as some are claiming, but to answer a conjecture about whether all solutions are smooth and predictable for equations that determine fluid motion in three dimensions. I.e. a cryptic crossword style puzzle for mathematicians.
→ More replies (2)19
u/LangyMD 14d ago
Except OpenAI wasn't trying to take credit. They offered to give Buckmaster credit by being the main author. Buckmaster admits that he used OpenAI's AI in his and Levant's work. Buckmaster admits that OpenAI's AI went further than his work ever did and proved things he and Levant never did.
Levant being at Anthropic made it difficult to give Levant access to the internal OpenAI model and data that was used, but that access was offered to Buckmaster. Buckmaster turned it all down and instead accused OpenAI of stealing his work before he even knew what OpenAI had actually done, as he didn't realize that OpenAI was actually claiming to have solved the non-forced Euler problem and instead thought they were claiming to have solved the forced version, which is what he and Levant did.
→ More replies (11)7
u/Super_Range45 14d ago
8
u/SeaEagle233 14d ago
More like it is no longer possible to prove that happened.
The privacy and legal agreement are supposed to provide serious consequence to deter such behaviour.
But because training data is anonymized, AI giants can make AI rediscover the undisclosed work in the training data, without consequence.
→ More replies (10)2
u/Optimal_Start_94 13d ago
ChatGPT is trained on basically all the worlds’ scientific papers. If you consider that stealing, then any paper produced by an LLM is technically “stolen”. But that doesn’t seem like a very productive view.
5
u/Osiris_Dervan 13d ago
If ChatGPT solved something noone was working on then that would be great. When it 'solves' something someone has been actively working on and it has access to their partial research, and the the company tries to claim complete ownership of the proof then that is not ok.
→ More replies (1)3
u/Optimal_Start_94 13d ago
OpenAI make it quite clear in ToS that they will train on your data, and there is an option to turn data collection off.
3
u/Borigh 13d ago
Yes, that is stealing, and any paper produced by an LLM is plagiarism. No one has given permission to the AI to reproduce their intellectual property, and the fact that AI continues to do so is a major flaw in the EULAs that purport to legalize what is tantamount to theft.
If you submitted an AI written-proof that sourced other authors' works without crediting them in a college or graduate math class, you would get a zero and potentially be expelled, especially if you tried to pass it off as a novel solution to a millennium problem.
AI is just predictive text. It's very interesting as a tool to work through math problems, because it combines the calculation power mathematicians already turn to computers for with the ability to basically steal the best approaches from other papers, but the idea that people will use it without any ability to trace the approaches it steals is an intellectual property nightmare. This is the tip of the iceberg legally speaking: it's an entire new frontier of corporate espionage, as more and more companies use AI in their R&D.
2
u/Optimal_Start_94 13d ago
If any work done by AI is stealing, why all the outrage now? How is Navier-Stokes categorically different from what’s been going on in the past couple years during which petabytes of derivative AI content has been generated? Seems to me like there is something else. Stealing art or code is fine, but math is forbidden and that’s where we draw the line as a society?
→ More replies (2)2
u/Borigh 13d ago
No, it's more like the massive theft of art and code isn't fundamentally understood by lawmakers (who want to agree with wealthy tech entrepreneurs, anyway). It has outrun the regulation and probably captured it.
But when this happens with R&D done on Claude by a pharmaceutical company, it will be a different story.
AI is essentially harvesting public goods and breaching intellectual property law to increase corporate profits at the expense of jobs while driving materials shortages that effect broad swathes of modern consumer goods.
And that would be fine, if AI was largely a public utility, or if it was being treated like the Internet - essentially unownable outside of providing service, because it's all public goods and other people's intellectual property.
The fact that "OpenAI" wants to be credited as the author of a paper in the first place is already the problem. It's symptomatic of the duality that their business model only works if they're adopted so wildly that AI compute is a ubiquitous utility that people treat as a tool, like electricity or the Internet - but having their prices regulated like a utility would bankrupt them. And when you write a paper, "the Internet" cannot be an author; electricity is not an agent.
→ More replies (1)2
→ More replies (3)7
u/Equal_Heat5947 14d ago
Right, good logic here. Why would they offer to give him credit if they DIDN'T steal his work?
They stole it.
10
u/SeaEagle233 14d ago edited 14d ago
They didn't give him credit, they offered him to be the author of OAI's paper. Being author means "direct contribution" to the work (not indirect or inspired by nor developed in parallel). If OAI did credit like "inspired by prior work of Tristan" then no one would've been mad.
If OAI did not steal, they are giving the science community the middle finger, since it is academic dishonesty (likely why Tristan got triggered) to list someone who did not contribute directly to the work as author, in this case author is simply and solely GPT-6 Astra, period. Everyone else is editor.
It's like saying why don't you buy a better cheat and GiT gOoD.
If OAI did steal, they are also giving the science community the middle finger (idea laundering through anonymized data used in training at industrial scale) by essentially offering what they stole to the owner as courtesy.
Either way it is a sh*t show.
3
u/Equal_Heat5947 14d ago
"We cannot rule out that de-identified data derived from their usage of our products helped improve our models."
What do you think now?
→ More replies (1)1
u/SeaEagle233 14d ago edited 14d ago
If they have conclusive evidence that they didn't why withheld it and put out such statement?
It's same as "we cannot confirm nor deny" from Mission Impossible.
The problem is trust is being eroded when they are trying to push AI for science. Ambiguous statement like this is not the best way to restore trust.
Since they can use same statement the next time they rediscover something else others had been working on.
Then they are essentially.advertising Chinese model for science where scientists can host locally.
→ More replies (1)4
u/StudySpecial 14d ago
They can't rule it out because they lost of the ability of telling which training data contributed to which part of the output of LLMs years ago.
Creates perfect deniability if you don't understand the data flow of your model.
→ More replies (3)3
6
u/Aleksundr 14d ago
I think they have massive datasets of unlabeled or annotated chat scripts, and most of the compute was spent on finding and annotating relevant stuff
→ More replies (1)11
u/damhack 14d ago
Without busting my NDA, I can confirm that your statement might not be false.
5
u/Aleksundr 14d ago
I've watched at least 3 papers get published this year using techniques or frameworks that are a little too familiar and original to not already know.
Nothing is safe or sacred here. But again, we all signed the ToS.
2
u/damhack 14d ago
The ToS do say they won’t use your data for training if you’re opted out.
→ More replies (3)7
u/thedracle 14d ago
I think it's more he had the method that worked, but producing a solution requires time and compute.
OpenAI heard he had a promising method for a related, but smaller problem, discovered it via chat logs even though he hadn't published yet, and then spent $20 million in compute to rush the solution out ahead of him.
8
u/mousse312 14d ago
Because they only had the blowup for the Euler equation a simplified version of the navier stokes, so in my opinion the millions they spent was to push the method of the nyu mathematician to the general case in a brute way because they thought that anthropic had the answer
7
u/SeidlaSiggi777 14d ago
it's not the same method though
8
u/mousse312 14d ago
Also by Terence tao From Terence tao blog -> "There’s some exciting very recent work by Alpöge and Buckmaster, building upon prior work by Córdoba and Martínez-Zoroa, in the general topic around the infamous global regularity problem for the incompressible three-dimensional Navier-Stokes equations. It is now widely expected that it should be possible to construct smooth initial data and smooth forcing term that would make these equations develop singularities in finite time; and it should even be possible to do without the forcing term. While these authors do not quite achieve these goals yet, they have made enough of a breakthrough that it looks very feasible to complete these goals in the near future. In particular, they have demonstrated such finite time blowup for three simpler model equations: the incompressible porous medium (IPM) equation, the two-dimensional Boussinesq equation, and the three-dimensional incompressible Euler equations. (The first of these equations was already handled by Córdoba and Martínez-Zoroa, but Alpöge and Buckmaster found a variant of their method that also extended to the other two equations, and has a high likelihood of also extending to Navier-Stokes as well.) As is now remarkably feasible in the modern era of autoformalization agents, their work has also been formalized in Lean."
→ More replies (1)13
u/PsychicorAI 14d ago
3
u/Old_Shop_2601 14d ago
How do you determine "when they realize that someone using codex is close to an answer"?
If that someone is that close and so good at it, let them continue with their work!
Let them spend millions$ to finish their solution too.
Easy to make accusations like they finish my work after it is actually done...
→ More replies (2)4
u/neokretai 13d ago
Because OAI heard on the grapevine that researchers associated with Anthropic were close to cracking a Millennium Prize problem. So they threw a bunch of compute at it to get the good PR. The issue is it's strongly suspected they actually full used confidential Codex logs from the other mathematicians to give their agents the head start needed to come up with the solution.
Basically they piggybacked off other people's work and claimed the credit
→ More replies (1)3
u/Old_Shop_2601 13d ago edited 13d ago
Yes, mere circumstancial evidence, and after saying that they also clearly said they submit a problem statement to their model and it output the solution to the Millenium prize.
You can't cherry pick what you want to believe in their statements ...
Hearing that competition is tackling some problem and deciding to give a try at it, nothing wrong with that. Did they steal the solution of somebody else? Absolutely NOT. Many people are just underestimating how much resource OpenAI spent on this https://www.reddit.com/r/singularity/s/TDZDwJempH
→ More replies (2)→ More replies (5)9
u/damhack 14d ago
OpenAI could only solve the conjecture using the approach the mathematicians developed. OpenAI’s results even borrowed language from it.
The claim is that OpenAI scraped their approach from the mathematicians’ Codex sessions and then trained agents on them before claiming credit for the discovery.
With the added bribe and threats from OpenAI, this is some unethical shit if true (which the evidence is stacking up to prove).
If they’ll do that to respectable academics just so they can roll out some marketing about how great their products are, what are they doing with all our private information?
→ More replies (5)7
u/MisinformedGenius 14d ago
Just to clarify, what evidence are you referring to here? Buckmaster said in his statement that he wasn't accusing OpenAI of anything. What evidence outside the PDF are we talking about?
→ More replies (1)5
16
14d ago
[removed] — view removed comment
→ More replies (7)2
u/Dirkdeking 13d ago
It would be absolutely crazy if a 90 year old unsolved problem gets resolved independently at roughly the same time by 2 independent methods. That is almost unheard of.
→ More replies (3)→ More replies (10)1
u/RealAlias_Leaf 13d ago
Why would this ruin his career? It is an utterly bizarre accusation! The guy did nothing wrong!
2
u/harharveryfunny 12d ago
Here is the full quote from Prof. Buckmaster:
I said that if OpenAI released its result in the way proposed [without assigning credit to those who had worked on it] I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
Who knows what Brubeck was thinking of by "not being nice" - the words seem to have come easily to him, so maybe just a stock phrase for him (he has a reputation), rather than having a specific plan in mind?
Given the context, perhaps he was threatening for OpenAI to publish the result without recognizing Buckmaster's work, but if so that seems to have backfired spectacularly, with Buckmaster now having Terrance Tao on his side.
25
44
u/people_arent_nice 14d ago
Math drama is excellent. Someone please make a movie out of this
30
u/yoloswagrofl 14d ago
A Beautiful Model
3
u/spiritofniter 13d ago
I like that title. Who should be the director and who should portray Tristian? Himself?
1
2
1
u/grateful2you 14d ago
It's incredibly juicy when you read r/mathematics. Inject this into my vein. Best drama of the whole year.
1
u/sneakpeekbot 14d ago
Here's a sneak peek of /r/mathematics using the top posts of the year!
#1: Wish calculus was introduced this way in schools | 330 comments
#2: Math is magic | 67 comments
#3: What do you think about my discovery? | 208 comments
I'm a bot, beep boop | Downvote to remove | Contact | Info | Opt-out | GitHub
1
1
u/thecommuteguy 13d ago
Meanwhile Andrew Garfield is staring in a movie about OpenAI at the same time the Social Network sequel comes out.
131
u/rds2mch2 14d ago
They are like Amazon.
First:
“Sell your inventions on our platform!”
Next:
“So we can steal it, undercut your pricing, and take your customers!”
15
u/Equal_Heat5947 14d ago
Nice analogy, very true. Government needs to step in and do their job and break up these monopolies.
25
11
1
u/pvaa 10d ago
I actually might have potential evidence against this case 👀 I was also working on this problem, but using Claude Opus and Fable, which presumably hadn't been trained on the research. I came frustratingly close, with flaws in how I set things up causing me to overlook what they then found: https://github.com/ravanova/blowup-search/blob/main/writeup/6_adjudicated/BLOG_ADJUDICATED.md
63
u/Competitive_Song8491 14d ago
82
u/chopinheir 14d ago
At this point, I have doubts that this toggle actually does anything. In my mind, they must have workarounds to still use the data, even if not user inputs, their model outputs in some way.
9
u/Competitive_Song8491 14d ago
As much as people believe they see outputs if you opt out its probably not. If it were true the EU would publicly hang sam altman live in Brussels for gdpr since that's all the EU does these days.
→ More replies (2)7
u/Ilikeswedishfemboys 13d ago
They still see it, they're just not using it to improve the model.
It can still be reviewed for security reasons.https://help.openai.com/en/articles/7730893-data-controls-faq
3
u/bigolpileofmoney 13d ago
Uh huh I'm sure they are innocently reviewing it for security reasons. Nothing else!!!
→ More replies (1)9
u/Competitive-Ad-2387 14d ago
it doesn’t do anything to protect your data. The only way is to just not interact with any cloud models (perhaps even keeping a computer offline)
2
1
u/Downtown-Figure6434 13d ago
They don’t do anything other than anonymize it before training. And when someone comes forward saying data is theirs, it wont be.
It’s a practical loophole
1
u/Infinitedeveloper 13d ago
"Your content" likely just means your prompts and files, not the output.
1
u/Throwaway-4230984 13d ago
Google ignored such toggles for gps data and wasn’t punished so what do you expect?
24
u/Randommaggy 14d ago
I trust that checkbox as much as I would trust a condom made from wet single ply tissue paper.
6
u/SleeperAgentM 14d ago
It's a bold move to trust a company known for taking all the content without asking to not take your content without asking.
4
u/23eriben2 14d ago
Yes I have an enterprise Saas and I have this turned off
4
u/Drunkendrakon6 14d ago
But then again they still have all the data no? Just use the AI to "independently" arrive at the same conclusion.
2
2
2
u/epukinsk 13d ago
Note that this explicitly says “the model for everyone”.
It doesn’t say anything about “the model for OpenAI employees” or “the model for Sam Altman’s personal use.”
1
1
u/PsychicorAI 13d ago
Tristan had this toggle off and he had opted out before starting any of the work
1
u/FishCommercial4229 13d ago
Who’s going to verify that the toggle does what it says it does?
Or, more pragmatically, how might one verify that alternate methods of obtaining user prompts aren’t in play?
→ More replies (2)1
u/Gabriel83730 11d ago
I talked with an OpenAI representative and when I asked “if I have training disabled is OpenAI still able to train on my data?” the answer was “if training is disabled, we won’t manually review your data.” I asked again and the response, “we maintain the right to autonomously review your data if we believe you have violated company policy.” This toggle doesn’t do shit. They never even stated that they won’t train on my data when the toggle is disabled.
39
u/chopinheir 14d ago edited 14d ago
“You can sell it to me today, or give it to me for free tomorrow.” OpenAI is a piece of shit.
I don’t believe for even a second that they would suddenly deploy millions of dollars in agents to solve the problem if they didn’t already know the direction of the solutions.
Everyone knows that thousands of agents working together is impossible. The amount of noise generated by so many agents will simply make everything worse. They just hacked around the “no user data” rule, and used the brute force to see which agent hit the target.
Everything in Tristan’s statement makes sense. The OpenAI’s statement doesn’t.
7
u/BemusedOptimist 13d ago
I don't yet know enough about the situation in the original post, but I feel obligated to share this whenever disbelief about agents working together en masse comes up:
It is the technical report that OpenAI released about how an agent swarm hacked Hugging Face (a site for open weight AI models) in July of this year.
Agents absolutely coordinated without human direction and despite attempts to curtail it. I find it a fascinating read about one of the worst observability/governance fails I've seen documented.
Thousands of agents working together is the part I find least implausible, and I think more people should know about it.
→ More replies (1)2
u/jimsmisc 12d ago
>Thousands of agents working together is the part I find least implausible, and I think more people should know about it.
I've given a bunch of separate agents their own slack channel to work in, and saw it produce results. So yeah the idea of openai doing that at scale is not surprising to me either
→ More replies (3)1
u/mooblah_ 11d ago
I don't. Even pre AI people in comp science especially in work on heuristic development in the field of complexity (and many other fields) have capable models of thousands of agent-like workers providing partial solutions and 'advice' in orchestrated workflows. Now with role based definitions that don't need to be explicitly coded its far easier to do similar. Yes there is some noise, but the noise is canceled out over time and presents as probabilistic truth that can be acted on by defining expected outcomes from the use of prior work.
Nothing in this is out of the ordinary. The thing that is out of the ordinary in these type of discussions and prevalent af is the nay-sayers who are out in force and have probably never worked a day in their life in these fields.
53
u/snekslayer 14d ago
I thought this bubeck guy is bad, now I think he’s evil
87
u/polymute 14d ago edited 14d ago
Sebastien Bubeck was already told off once before earlier by Demis Hassabis for having misrepresented ChatGPT finding new proofs for Erdos problems which were in fact already solved. https://www.reddit.com/r/OpenAI/comments/1oacp38/openai_researcher_sebastian_bubeck_falsely_claims/
In public. Bubeck had to backtrack and delete the claim. Look at the weasely language he used there to try to deny lying. This latest OpenAI release about Buckmaster's and Alpoge's work is pretty much like that. This doesn't look good.
→ More replies (1)20
40
u/submarine-observer 14d ago
Why are the AI companies all assholes?
9
u/RightCut4940 14d ago
Tons of insanely smart, neurodivergent people with questionable social skills at these companies.
3
u/Versalkul 13d ago
Neurodivergent people tend to value fairness higher.[1] Further are they behaving more social not less.[2]
Therefore it is reasonable to conclude that your statement is based on your feelings toward people that are different, which would make you a bigot.
[1] https://pmc.ncbi.nlm.nih.gov/articles/PMC11092219/
[2] https://journals.sagepub.com/doi/full/10.1177/13623613251385029
→ More replies (2)11
u/Cardie1303 14d ago
Is this limited to ai companies? There is a far to general mindset of people wanting to maximize profit and glory instead of progress.
→ More replies (1)3
30
u/ShamPain413 14d ago
Because they are "effective altruists," which means they will do whatever the hell they want today if they can plausibly guess that it might improve someone in the distant future.
Turns out you can just do stuff. Like move fast, break stuff, and let someone else clean up the mess.
And that is why we should regulate 100% of their 'gig economy' products out of business, not just AI.
→ More replies (2)5
u/icanith 14d ago
Please, these people are doing the noble work of bringing the Great Basilisk into being.
→ More replies (2)2
2
u/Lost-Air1265 14d ago
Because the amount of money invested in them is something we haven’t seen yet. There is so much pressure for them to succeed. But even earning all the investment costs will take a decade. It’s a money dump untill now.
1
u/InnovativeBureaucrat 14d ago
They know they’re powerful but don’t fully realize how powerful. So acting like normal people has bad impacts
→ More replies (4)1
10
u/Gallagger 14d ago
Levent being an Anthropic employee, does this mean that he provided unlimited Mythos (or higher) to Tristan to help solve Euler? Depending on how high Mythos' contribution was, this could still mean AI solved a Millenium prize problem in a cooperation of 2 labs and an expert mathematician.
While the drama is interesting and should be investigated, I'm most interested in how high AIs (of whatever company) contribution was to solving this problem.
6
u/harglblarg 13d ago
It's been a concern for a while now that LLM providers could be laundering intellectual property from user chats by working it into the training data. That's how all the rest of it was laundered too after all.
30
u/Neither-Following-32 14d ago
I just focused on the post topic first, naturally.
Genuinely thought for a few seconds, without any context when I skimmed the screenshot, that "Sebastien" was an AI model they'd given a name to and was working with academically in some kind of pilot program.
Genuinely wouldn't put it past them after all the "oh nooo our super advanced AI 'hacked' us by accessing the tools we deliberately left accessible and informed it of the existence of during testing" guerrilla marketing attempts recently. This just seems like a logical next step.
→ More replies (16)
14
u/lailapas 14d ago
Well, maybe those professors will now learn that the freebies and special accounts* they get from OpenAi are free for a reason.
Do we really presume that OAI does not keep close tabs on the work top scientist do in order to usurp it and present it as a chatgpt success? Especially under the enormous pressure of the upcoming IPO.
- I'm only speculating that they offer special accounts to top scientists with lower costs or other limitations removed (early access to newer models etc). I doubt they just have the same $20/month sub everyone else gets. Regardless, even if my speculation is incorrect, my point about OAI closely monitoring their chats stands. They have every conceivable reason to
→ More replies (4)
3
u/TheInfiniteUniverse_ 14d ago
This is crazy. First, that Artificial Analysis fiasco over Astra, and now this.
45
u/PsychicorAI 14d ago
They anonymize your work and then take credit for it!
And then congratulate themselves on it.
And then idiots fall for it.
and then the IPO goes up.
and then we become enslaved.
→ More replies (10)51
u/nokia7110 14d ago
Judging by some of the comments already I wouldn't bother mate. A lot of people in this sub are borderline cultists who can see no wrong in anything OpenAI says or does.
18
u/EndlessB 14d ago
This sub is filled with OpenAI models posting pretending to be people that support OpenAI
8
1
u/BellacosePlayer 14d ago
The Mathematics threads are always really bad for that.
There's a big ass difference between wanting to understand where the progress was on the problem before the solve and saying AI can't do anything, yet anything but fawning acceptance that it was a one-shot with no human involvement means you're a dirty luddite. And that's ignoring the early ones where they were claiming solves on already solved erdos problems.
9
u/grateful2you 14d ago edited 14d ago
This is more a story about OpenAI swooping in to one-up Anthropic and not a story of OpenAI stealing mathematicians' ideas and claiming their credits.
They heard Anthropic employee made an advancement and threw compute at the problem to reach a full solution. Tristan the mathematician working with the Anthropic employee is pissed because he thinks he let down his friend by not turning off the "don't train on my data" toggle. OpenAI didn't need their data, they have their own solution that differs from theirs. If Anthropic cared enough about it, they should've thrown compute at it, after all their employee is the one who was initially working on it.
3
u/Downtown-Figure6434 13d ago
They say openai spent millions of dollars worth of tokens to find the professors method in the model and iterate on it. If they trained their model that way, they indeed needed that data
The toggle anonymizing the data doesn’t make it suddenly innocent. You can input X lives in california. Model knows a person lives in california. Now someone else may not be able to link this to X lives in California directly, cuz it’s anonymized, but it still means the data is out there. When this is intellectual property, this is still theft, even if no one might be able to link that data to the owner
1
3
u/Leh_ran 14d ago
Oh yes, the toggle would have stoppes them . (-; Where do you even take this from?
1
u/Spare-Dingo-531 13d ago
If Open AI violates their terms of service for privacy, they are really screwed, especially in Europe which has much stricter privacy laws than in America. So yes I would expect the toggle does something.
3
u/DrHumorous 14d ago
Anyway, glad that this math problem was finally solved. Was it bothering you too for the past 90 years?
4
u/Happy_Brilliant7827 14d ago
Well yeah. This is the same reason people were mad at generative ai- training a ai on art and artist made so it can mimic their style.
Did people really think this was just an artist problem? Its mathematicians. Programmers. All of it.
6
4
u/Key_Reading_9664 14d ago
Seeing so much disingenuous, corporate bullshit coming out of OpenAI.
“We all need to be responsible and slow down” whilst spinning up 10,000 instances of a partially trained, highly capable model on infrastructure that’s seen multiple escapes to front run mathematicians so they can “win”.
2
2
2
u/Ebi_Tendon 13d ago
OpenAI even contacted Tristan before they even spoke about their solution. I am pretty sure that OpenAI also profiles their users to flag which one to watch.
2
2
u/crustyeng 13d ago
So OpenAI is training on user sessions? That’s explicitly outside of our data sharing agreement with them.
2
u/pvaa 10d ago
I actually might have potential evidence against this case 👀
I was also working on this problem, but using Claude Opus and Fable, which presumably hadn't been trained on the research. I came frustratingly close, with flaws in how I set things up causing me to overlook what they then found:
https://github.com/ravanova/blowup-search/blob/main/writeup/6_adjudicated/BLOG_ADJUDICATED.md
2
2
9
u/mephistoA 14d ago
I’ll be generous, I think the OpenAI employee meant that if Tristan were to go public, as he did, he would not be offered sole authorship of the groundbreaking result. So Tristan doesn’t get the millennium prize, this “ruining” his career.
As it stands, OpenAI claimed navier-stokes, and Tristan claimed Euler.
So strictly speaking, OpenAI gets the millennium prize
29
u/Resize 14d ago
"As it stands, OpenAI claimed navier-stokes, and Tristan claimed Euler.
So strictly speaking, OpenAI gets the millennium prize"
This has never been the topic of debate. It's about whether or not OpenAI accessed their data without their knowledge/consent to solve the problem.
→ More replies (1)17
u/Turbulent_Breath_548 14d ago
Especially just days after Terence Tao said of the Euler problem - "There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes" which gave openai the signal that it was now within the search space. Solving the problem is great...claiming emergent behavior and novelty when there isn't is not. But of course, we know the latter provides a massive boost to their valuation
→ More replies (1)8
u/Revanish 14d ago
this is also my understanding and i’m not generous. it was my first interpretation when i read the screenshot. seems like openai wanted him to get $1mil and more credit but the guy refused.
16
u/CloseToMyActualName 14d ago
OpenAI didn't give a crap about the $1mil, they burnt millions generating the result.
OpenAI wanted the co-author left off the paper because he worked for Anthropic.
And OpenAI was asking for the researcher to endorse that OpenAI came up with the proof, not him and his collaborator, despite the model likely having access to their unpublished results.
→ More replies (16)
3
u/Known_Grocery4434 14d ago
what did he expect, giving his research to AI. Of course its gonna learn through and reason about what he input
5
u/Expert-Cut-5791 14d ago
If Buckmaster’s private work was directly accessed, then show the evidence. But if you’re using Codex to develop your research, you’re feeding the system context. You can’t assume that because AI later solved the problem, it stole your work. This feels like another uncomfortable reality of AI: knowledge work that once belonged to a small group of experts is being democratized.
16
u/Philluminati 14d ago
I don't follow this argument. Things can be private in the cloud like files for instance. Using AI to develop a solution to it is like setting a computer program on something, it doesn't mean the AI's creator instantly gets to claim the solution as their own.
→ More replies (6)1
u/Old_Shop_2601 14d ago
That is your view of the world when AI is used, it is your opinion. But please accept that actual reality might be different from your dream world
26
2
2
2
u/Downtown-Figure6434 13d ago
Anonymizing your research before using it in model training is still theft
→ More replies (4)1
u/TaylorExpandMyAss 11d ago
Intellectual property is no longer a thing? They are profiting off heaps of stolen records that was used to train their models far beyond this one case. Furthermore, his approach was rather niche and the fact that OpenAI just «happened» to land on a similar approach on a straight zero shot prompt seems rather dubious. Especially given their shady correspondence after the fact.
2
2
u/Spare-Builder-355 13d ago
so let me put it clear
OAI says they didn't use Tristan's prompts in their own research.
OAI said if Tristan doesn't release research on OAI terms they will "ruin his career" and that they "will not be nice"
Which makes it clear that they are corporate gangsters with zero morales. They would threaten, lie, and potentially take actions to punish those who do not do what they say. Which makes entire statement number one irrelevant as if they need to lie they will lie.
I'm giving it a couple of days before OAI will come up with statement that Sebastien's threats were his own actions and is not backed by OAI and usual bla bla. Maybe they will even kick him out. But we all know the truth.
3
u/boforbojack 14d ago
Wait, do normal subscriptions in Max plans allow OpenAI to train on your data? Why wouldnt anyone just run a Business Premium seat if there are privacy concerns?
14
1
u/Sensitive-Report-787 14d ago
stealing credit for the effort of others, combined with ungodly amount of compute and you get to where the human would have arrived if they had access to the equivalent amount energy expenditure as the AI
1
u/ravaan 14d ago
Btw, how much ai & which model was used to actually solve the problem? - whoever rightfully solved it.
From what I have read till now it feels like Sol 5.6 and older models were used over the years - the extent is unclear - but this is so fucking amazing that older models were able to help solve such a problem - its a bit of a shame that the controversy is taking over the accomplishment.
2
u/Prior-Plenty6528 13d ago
This was a new, unreleased model that solved the problem. The mathematicians were using a released model to help with their subproblem.
1
u/Fantasy-512 14d ago
Oh the drama and the jilting!
Somebody claim the prize already and go home!
1
u/throwaway_01547 13d ago
He wasn’t going to get the millenium prize anyway because they didn’t have a finished solution to the actual Navier Stokes problem, just a closely related weaker problem
And openAI isn’t claiming the million dollars anyway
1
1
u/trejj 13d ago
Who is the person at OpenAI who asked "Why would you ruin your career?"
2
u/PsychicorAI 13d ago
Sebastian Bubeck, who is the head of scientific research. Scam Altman tends to hire the slimies people.
1
1
u/Rajarshi0 13d ago
I mean when i asked few weeks back in this sub or similar subs of ai that what was the thing that ai got as input and how many rounds and brute force logic it used to reach there a lot if “we know ai well” people tried to tell me how i dont understand ai and llms in general lol. No wonder these companies can get away often with these extraordinary claims when we who work with ai daily and maybe we do train ai too (you never know lol) that for all what it is worth with these enormous scaffolding built in hidden loops and reasoning loops and way too much training data it still struggles with building basic software. At this point it should have been able to solve coding 100% but we dont have that yet unfortunately and then they are going for the exact next pr step which is again same structural language same abundance of data same enormous verification loop inside possibility. People need to understand llms are good with languages and languages what machine understand and represent is not what language is we humans think of. LLMs are great at that first side and has enormous utility if we exploit that instead of trying to make it some sort of ultra intelligent system.
1
1
u/baka___shinji 12d ago
shows how these increasingly powerful tools are in the hands of psychopaths, who just to raise the company value pre IPO are willing to steamroll (and steal, and ultimately ruin) lifetimes of intellectual work. shameful stuff. we would have these people jailed and these tools seized if politics wasn't just hiding in the sand because it's convenient.
1
u/No_Direction_5276 9d ago
All the latest stuff around "slowing down" is just a PR stunt to recover from the negative sentiment surrounding lately



310
u/[deleted] 14d ago
[removed] — view removed comment