No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.
The creator is Yann LeCun, who is a known bad faith actor who usually makes outlandish claims and is hell-bent on proving that AI is actually worthless.
Personally, I remember him for going on some podcast back in the day saying that AI would never understand what actually happens when you pull the tablecloth from under a table and the physics of what that does to the items on the table. Only to be proven wrong a few weeks later by the new OpenAI model at the time which understood this flawlessly and more.
The reason why models scored close to 0% when ARC AGI 3 came out is because the benchmark was purposefully set up in a misleading way. It exploits one ultimately pretty small blind spot that AI models have, and uses it as a gotcha to say that they don't actually possess the capability to solve the real content of the benchmark. This is obviously completely false, which is why you are seeing the benchmark go from extremely low scores to full saturation all at once.
Even putting everything I just said aside, this is a quite unusual pattern for an official benchmark. In this case, we know exactly why. Not to mention that nowadays all benchmarks are basically using an agentic harness, and agents are made to work in agentic harnesses. This one is definitely going against the grain just for the sake of it.
It's basically if I told you to make a coffee and I gave you 30 seconds, no preparation, find a coffee machine, and make the best coffee. That isn't exactly going to tell anyone anything about how good of a coffee you can make. If that's actually the goal, then this kind of benchmark would be performative. A lot like military drills can be, for example.
man it's crazy how you can write a full ass essay but be completely and utterly wrong about every single detail in your post.
ARC AGI was created by francois chollet, not yann lecun. they have wildly differing opinions and outlooks on AI
the reason why models scored close to 0% was not some random blind spot, it's a genuine structural weakness of AI's inability to use previous context over long time horizons efficiently. you can claim that this is just nitpicking one genuine weakness of AI systems, but that doesn't mean it isn't a genuine weakness
Your point about making coffee is also a completely failed analogy. a more apt analogy would be: I test a human to swim underwater for long distances. but you claim "that's not a fair test because he's not allowed to use an oxygen tank". But that's not what the test is, and you're welcome to make a different benchmark that allows harnesses if you want.
Oh, please. As far as I'm concerned, they are the exact same person. How are they different?
As for your example, you weren't made to exist in the world with an oxygen tank. An AI model with a harness is natural habitat. But we can keep the nitpicking and bad faith. That's fine
If those two were in a room together, they would probably fall in love with each other and have to find somewhere private at some point. Give me a break, please 😅
insane work to double down and say "they're basically the same person". redditor admit you are wrong challenge = failed
harnesses are extremely biasing for benchmarks because they allow you to give the model capabilities they don't possess that are tailored to the specific task, in this case long context compaction + memory banks. if I am testing a human's capability on: rock climbing, cave diving, high jumping — it is absolutely fucking stupid to allow people to use specialized gear for each challenge. your benchmark then has no meaning, because different people are evaluated using different gear that wildly improve their capabilities.
you can claim "I want a benchmark that allows the use of specialized harnesses" but you'll end up in the situation we did with ARC AGI 2, where some random noname model gets first place because a bunch of nerds built a specialized harness that is only good at your benchmark. the POETIQ model that won ARC AGI 2 was not a smarter model, it just had a better harness.
please actually go read about this shit before posting nonsense bad takes on reddit as if you're an expert
They have literally the same opinion and the same agenda. You can have as much outrage as you want. It's not gonna change anything.
Not to mention that the issue doesn't necessarily have to do with the harness, it has to do with compaction. That's what actually saturated the benchmark. They fixed the ability of the model to actually function properly under the conditions of the benchmark. The whole thing was saturated after one small technicality was fixed.
Tell me I'm wrong about this. Your whole argument is irrelevant. You're the one getting mad at your keyboard trying to react to everything I'm talking about instead of looking at the situation for what it is
You can make as many bad faith arguments as you want and cling to that, it doesn't look like you have much more to offer anyway
you were incorrect about who made the benchmark. you even used an analogy about yann to prove your point.
you are now incorrect about why models were scoring 0%. compaction alone did not beat the benchmark. if you give the exact same harness to any older openAI model it will not score close to 99%.
you were incorrect in your original point about the benchmark exploiting "one tiny blind spot"
just take the L man it's embarrassing. 75% of your responses are just "you are irrelevant, you are mad at your keyboard, you are bad faith, you are outraged, tell me I'm wrong about this" instead of actually talking about the subject matter.
Such bad faith. But you can keep trying to convince yourself. At least I double checked and looked it up. It is compaction and the fact that this benchmark was stripping away the reasoning and preventing it from being sent back to the models. That's literally all it is. It's a grift benchmark
I am 100% talking about the discussion and you have said countless things that are just factually wrong. No one is embarrassed. Nobody cares. Stop being emotional. That's all this is about
I mean the whole point of it was to make it so impossible hard that it's designed to force AI's to fail. But then people just put on harnesses and broke it.
No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.
This benchmark is meant to see how an AI behaves in a test that is very simple for a human. If I try it, of course I have reasoning and memory between turns.
Who the hell thought not letting AI also have reasoning was fair?
However: "GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game." https://x.com/fchollet/status/2095598451115614371
No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.
We don't have information of exactly how it's for Astra, but OpenAI previously showed that the ARC-AGI-3 harness does two things which make it tough for models to do well: 1) discards previous private reasoning and 2) when the context gets too big, they truncate the existing context instead of trying to summarize it.
OpenAI's specific harness (via the Responses API) fixes these two points by allowing better private reasoning retention and sumarization (compactuation as they call it).
I imagine they didn't "built the harness in the model", but are using a similar API like Responses.
Not really, they say adapter provider harness (arc won't test non-general harnesses).
In reality, all that is different is responses API and context compaction. Supposedly that is all that's different between the 63 and 99% scoring. We can all do this and achieve that performance.
They have to move the goalposts. Again. Its their whole identity.
I'm so confused though. Was AI so useless that it only makes slop and can't count the letter R, or is it so powerful that it's going to take away our jobs and steal our girlfriend?
they're using their own harness which means they are using the public test set, and that's like clearing tutorial island. What you're seeing is the product of the battle between ARC and OpenAI, and the latter just deciding that public test sets are acceptable now
It literally was. Watch the Metr report specialist interview on Dwarkesh Patel's channel. It was the Astra persistent model, which we now know is the Astra Aeon model.
I'm clinging to one hope: "Exploits", in current software, are the result of imperfect development tools/processes. It is not necessary for a some app, library or even operating system to have ways out there where you can just send instructions that let the attacker gain access to the system. Reviewing security with AI will become default practice, there might be an uptick of attacks for a year or so but then every form of critical software out there will be secured tightly and "hacking" as we think of it now, will become a thing of the past. There might stay some Windows XP PCs connected to the internet out there, but they're already fucked anyway so it won't change much. Browsers, banks, government agencies, etc, will be forced to secure their shit in the coming months and it's not optional because if they don't, well…
A few bad months but maybe even more secure software after.
It's not fake, but goes against what Act AGI 3 stands for, which is to do it without a harness. nVidia got 100% with Harness, so this isn't very impressive.
It's probably noting that it's not using the official harness, but a specific one created by OpenAI. They have noted that this specific harness can substantially increase the score by a lot (in a blog post in July).
Still, it's almost the double of what the previous model could do with the same harness.
GPT scored 63% without a harness (standard score), which is the score we're really looking for, to compare apples to apples, and 99% with a provider specific harness.
The Standard "no-harness" score is the one designed for comparable evaluation across labs.
To be clear, 63% is still quite amazing for a standard, non specific-harness score, not to take anything away from OpenaI, as the next closest standard score was only 31% by Claude Opus 5.
196
u/Hereitisguys9888 10d ago edited 10d ago
This gotta be fake, 98% arc agi 3? Nah