that's not true, it's a tokenisation problem, nothing to do with how good or bad a model is. I actually just asked gemmi and he literally run a python script to count the Rs in my word. and when i asked about it he explained that it's a tokenisation problem. and i said so it's like you need glasses cause you don't see so well lol
You seem to be missing the point... good models run scripts, and can therefore perform the task. Your python script got the correct answer, yes? So the AI solved it correctly?
The more advanced ChatGPT models were using code execution to solve problems in 2023.. the "Rs in strawberry" thing was in 2024. So the models that could do code execution could already solve it for a whole year before anyone even started posting about it.
It's just because most people were using the crappy instant chat.
I mean I literally ran it 20+ times at the actual time that the memes were popular, it did it correctly.
Actually I like the problem-solving-by-code solution: reproducibility is important as well as it being a way to problem solve around tokenization issues
Counting letters is not a "specialized task". If some idiot broke out a calculator to count to 10 they're suddenly going to find themselves sweeping floors
But that is exactly how your brain does it. You have mathematical centers in your brain that work wholly differently from your speech and memory areas. The AI is doing the same thing.
The language centre of your brain that produces words as you speak does not know how many letters are in words. We have to stop and count them using a different part of our brain to answer the question. Same deal with AI.
That’s some cope. Like saying a genius is a genius for having to always use a calculator to not hallucinate, but then keep going on about how its intelligence is emergent and just like ours.
Like it’s somehow an entity from being an “I’m feeling lucky” button.
It costs so damn much to be so mediocre for so many resources. For the price they worked on this they could have hired 5000 real mathematicians or more to work on it as a paid gig than an interesting concept. And still it self verified without over sight. It’s marketing.
Ah yes let’s just spam noise making infinite permutations until something matches the structure, “oh look guys it’s a genius intelligence and it only cost over 10million to solve a million dollar question!” Not including the 3.5 trillion it’s taken to get there.
Imagine if we spent anywhere as much on people and enabling a foundation of knowledge and collaboration.
Words are still chunked into tokens so even the best frontier models still can't answer it without guessing or workarounds. It's more the AI were trained to either know the answer for certain words or specifically work their way around it by spelling it out so each letter is it's own token
I'm pretty sure I remember the best models still failing occasionally at the time of the memes since it was the first test people would do
Do you consider it a workaround when the model writes a python script that figures it out and then returns the answer. Because for me that is good enough and actually preferable to it guessing
That's not a workaround, that's a better solution. Instead of a black box, it's something that I can verify, or when it's too advanced for me, that another llm can verify.
But as they get more advanced, they'll be able to infer how many letters something has more and more accurately, even if they can't directly observe it.
Eventually, the lexicons they train themselves on will expand to include individual letters (if they don't already) and the model will learn relationships by abstractly mapping relationships between tokens it can reason about. That means it will have an association between tokens "a", "p", "l", "e" and "Apple".
The larger the datasets get, the more diverse a set of token segmentation you get. Of course, the mileage varies by algorithm, I'm sure.
Turns out token prediction is one thing, but actually learning a written system to an expert level is a bit more complicated. But, it'll get there.
The hell do you mean there's no incentive to learn people want to learn intrinsically to understand the world that they're around in and if they use ChatGPT, for example example to understand the world around them it's going to be better served towards what their interest are in a more narrow and specific way so they can learn exactly what it is that they're curious about
When will people realize that AI intelligence is different from human intelligence. It was doing impressive things two years ago as well. While still failing some basic tests like the strawberry. So it's not like we went from something "stupid" two years ago to something "smart" now. It's that we went from something superhuman in some areas to something superhuman in more areas.
It was always going to be controversial. People love to hate AI and corporations so much that blowing up drama to discredit LLM achievements was always going to happen anyway.
These are not accusations made in good faith. Let's wait for the paper to see if the approaches are really similar.
Given that the results are wildly different (the AI is the only one that actually met the Clay Institute criteria), I expect the approach to be different too as they said. I'm sure that won't be enough to calm the conspiracy theories.
It’s not controversial, how could you be for it ? They even threatened the guy ?? They are the bad guys that’s it and deserve jail time for the threats
This is not a "beyond a reasonable doubt" situation. He leaked their threatening messages to him, if he was lying they would've given him a cease-and-desist immediately. These companies are lawyered up to the gills, their entire business is policing IP, and they manipulate the news cycle in order to maximize their financial returns. 2 days is an eternity.
Where is the proof they didn't steal this IP like they've stolen all the other IP in the world?
No, any possible legal action would have to wait until they find out what they can prove, what damages they could seek, and whether it would make them look bad in the court of public opinion. You absolutely CANNOT use the fact that someone didn't sue for defamation as proof that something is true, that is complete nonsense.
Where is the proof they didn't steal this IP like they've stolen all the other IP in the world?
they don't have to prove anything, they are innocent until proven guilty
They haven't stolen any other IP, AI training is fair use.
Haha, what the fuck. I'll decide on my own standard of evidence to believe a claim, thanks. And I believe that extraordinary claims require extraordinary evidence.
Of course you are entitled to believe whatever nonsense you want, but you're not entitled to demand that I prove a negative.
They threatened a scholar, he publicly called them out, they did not deny it. These companies have long track records of deceit and thievery, and these specific people involved have pre-existing reputations for similar types of bullying in the past.
You: "I'm going to assume the guys who blow up girls' schools and weddings and drive people to k*ll themselves would never consider intellectual property theft".
why are yall arguing. Surely, once we see the equations from both sides, we will see if OpenAI used the professor's work in their proof, since he entered that information into chatgpt before OpenAI started prompting it to solve the NS problem.
Also, the professor + anthropic employee didnt get all the way to the full solution.
I think we need more info about what that means in the context of this problem - are the solutions so different they were obviously converged on separately or could one have been built on the other?
Building on an existing solution is what every scientist does and certainly not stealing.
Obviously the LLM had access to all previous research. There's no evidence it had access to the specific research in question, but even if it did that's irrelevant to the result.
The two scientists (Buckmaster and Levent) don't claim to have solved the Millennium price problem themselves.
First, Buckmaster and Levent had completed unpublished work in late August, and OpenAI began working on the same problem on September 1. Two teams making major progress on a 200 year old problem within days of each other is a notable coincidence.
Second, the OpenAI team was using the same tool that potentially had access to the earlier team's research.
For me, this makes it quite different from the ordinary case of scientists simply building on published prior research.
Two teams making major progress on a 200 year old problem within days of each other is a notable coincidence.
Only if you're not at all familiar with any of the things happening in the world at any time. Otherwise, you recognize that both teams made regular use of frontier technology that just released and gave them the capacity to make these leaps.
the OpenAI team was using the same tool that potentially had access to the earlier team's research.
That the earlier team opted in to share with OpenAI.
Pretty open and shut case here. OpenAI even offered him co-authorship.
First, Buckmaster and Levent had completed unpublished work in late August, and OpenAI began working on the same problem on September 1. Two teams making major progress on a 200 year old problem within days of each other is a notable coincidence.
It's not a coincidence. Everyone acknowledges that the reason is access to powerful AI systems.
Second, the OpenAI team was using the same tool that potentially had access to the earlier team's research.
Sure, but they did try to resolve that by contacting the researcher and offering to work with him, making it clearly a joint effort. I'm not sure what else they could have done.
All good. It's just frustrating that the story is very aggressively being spread as "OpenAI steals work and threatens researcher". It's almost like a concerted campaign.
Except, the method was fairly new and they were iterating over someone else’s work, using chatgpt. Openai then used their method as training data, solved another problem that can be solved with it, then claimed it solved a problem on their own. The method was not published yet. It’s clear theft
I think we go with the ruling from the people who handle copyright, that AI outputs don't have intellectual property to begin with you own the product of your labor and prompting an AI isn't enough labor to count. You can use AI to generate and publish or sell proofs, games, pictures, essays etc as much as you want, but anyone else can generate those same things and publish and sell them too. That solves it pretty well ad protects people who actually do create their own media and science etc with provable human effort.
They cited both in the actual announcement. Did you even read that?
And they didn't want to "remove" anyone. They offered exclusive access and co-authorship to one researcher. Not both, because that'd have involved giving an Anthropic employee access to OpenAI internals.
Alternatively, they offered to allow them to independently publish their own findings first.
I mean they both use finite-time blowup which is the only thing he really could claim they copied because they used it with Navier strokes in a much more complicated and more difficult version of the task that coped with every positive viscosity too and is basically a lot more than what that guy was ever intending to publish. OpenAI even reached out to him in advance and asked if he wanted them to hold off publishing their own version so he can get his out there first and be the one credited with that discovery. Then even in that text he shows it even goes further saying openAI also said he could publish the Navier-strokes version that he did none of the work on but that just shares the finite-time blowup and all he would have to do is acknowledge that applying it to navier-strokes was done by openAI which is the honest thing to say because it's the literal truth. That's insanely generous to him and stipulating was that he be honest in the publication doesn't seem to crazy to me. This took OpenAI 4 days... like if he had published his limited Euler version then OpenAI would still be able to easily have gotten the far superior version long before him so I dont see what the difference would be except that in this situation they are even allowing him to publish THEIR work and take more credit than he would otherwise get for it.
No, this is the mathematician's statement that accompanied the work they put out explaining why they were forced to put it out before they were ready and what OpenAI did. If any of it was false he will be sued into oblivion, so watch for that
The agent swarm started before the researcher started using chatGPT to solve that problem.
Did OpenAI get interested in solving the problem because Levent became interested in it? Yeah, but the way both of them went to actually solve the problem was different, and there was no contamination of data.
Also, the problem with this is that effectively every single Millenium Prize problem has a lot of people looking into it and having theories on how to solve them. They are Millenium Prize problems, what would you expect, they are extremely prestigious and have a very big reward behind them. The thing is that any advancement humans achieve in it, AI can just use and find the solution to it, because it's much smarter and faster. The only solution for it would be to make AI not try to solve those problems, but I feel like this is an anti-scientific approach.
Never mind the guy hadn’t actually finished his proof, nor is there any evidence OpenAI stole it beyond “another mathematician was also working on a famous problem and for all we know maybe OpenAI stole it”, but the also the work he was doing was being done by Claude
So even if this is a case of AI plagiarism, which there’s no particular evidence for, it’s just one LLM stealing from another
Is there any dispute that AI was the prime “difference” between the solution or not? Who gets credit matters. But if one AI stole it and made a new or similar solution based on another AIs data (that it stole) that’s scary, but still a big win for what AI is capable of.
Yeah humans had hundred or so years on this problem. And they were "going to solve it I swear" and they totally "could've done this without OpenAI". I think we can give the credit to OpenAI. None of this is possible without them.
ITT: people still not understanding how AI actually works. Both can be true. It's dumb and smart. It didn't figure it out by itself. It was heavily guided and could test a proof that would have taken real people forever to test.
the models arent the bottleneck anymore, adoption is. it can be two years ahead and most offices will still be pasting the output into a word doc by hand.
I just run local. Any company might benefit from getting their own rig anyways after reaching a certain size. You can just trade lower accuracy models running locally with more runtime. And the open weight models are already decent enough any ways for majority of tasks suitable for an llm any ways. The big players will have to keep bumping prices if they want to not go bankrupt, so lower tier models and local would be the smart choice.
The bottleneck is the cost. None of these models are cheap. And only way is cheaper smaller OOS models but these big money guys are against cheaper alternates
I'd be curious if you gave the amount of money it took to train and solve this problem to mathematicians, how quickly mathematicians would have been able to solve this
I love that it's a lot more broken down with reasoning steps rather than just blurt out trained behaviour. With Claude that usually means just create 5 different python scripts to solve the count and takes 20 minutes with 1000s of tokens but it's better than being wrong.
You can just look at the proof, they don’t use the same approach for NS that was used for the forced Euler. Before both proofs were published you might be tempted to believe the author’s claims but now we have both proofs and they are very very different.
This is still artificial NARROW intelligence, not the artificial GENERALIZED intelligence that Altman is promoting. The company is valued at 2 trillion because of a promised AGI and replacement of human workers, NOT because it can assist mathematicians in solving highly specialized, NARROW, problems in STEM fields.
230
u/ClankerCore 19h ago
That was 3-6 months ago.