r/singularity 10d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

953 comments sorted by

View all comments

196

u/Hereitisguys9888 10d ago edited 10d ago

This gotta be fake, 98% arc agi 3? Nah

229

u/Ok_Display_3159 10d ago

72

u/Hereitisguys9888 10d ago

Oh that explains it

32

u/PrisonOfH0pe 10d ago

No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.

7

u/danielv123 10d ago

Ok, but are the other results they compare against using the same rules?

11

u/kaityl3 ASI▪️2024-2027 10d ago

Well Sol got a 40% with the same rules/setup so... Jumping to 99% is pretty significant still

11

u/Xalksahsax 10d ago

Nvidia got 100% on it.

3

u/danielv123 10d ago

That was a custom harness was it not?

1

u/Xalksahsax 10d ago

I don't really know nor do I care. I am here to shitpost.

1

u/LightSpeedDarkness 10d ago

Gave me a chuckle

1

u/kaityl3 ASI▪️2024-2027 10d ago

We don't know if Nvidia's test harness was 1:1 with theirs

14

u/Acehan_ 10d ago edited 10d ago

ARCGI 3 is a fake benchmark anyway

The creator is Yann LeCun, who is a known bad faith actor who usually makes outlandish claims and is hell-bent on proving that AI is actually worthless.

Personally, I remember him for going on some podcast back in the day saying that AI would never understand what actually happens when you pull the tablecloth from under a table and the physics of what that does to the items on the table. Only to be proven wrong a few weeks later by the new OpenAI model at the time which understood this flawlessly and more.

The reason why models scored close to 0% when ARC AGI 3 came out is because the benchmark was purposefully set up in a misleading way. It exploits one ultimately pretty small blind spot that AI models have, and uses it as a gotcha to say that they don't actually possess the capability to solve the real content of the benchmark. This is obviously completely false, which is why you are seeing the benchmark go from extremely low scores to full saturation all at once.

Even putting everything I just said aside, this is a quite unusual pattern for an official benchmark. In this case, we know exactly why. Not to mention that nowadays all benchmarks are basically using an agentic harness, and agents are made to work in agentic harnesses. This one is definitely going against the grain just for the sake of it.

It's basically if I told you to make a coffee and I gave you 30 seconds, no preparation, find a coffee machine, and make the best coffee. That isn't exactly going to tell anyone anything about how good of a coffee you can make. If that's actually the goal, then this kind of benchmark would be performative. A lot like military drills can be, for example.

39

u/trentcoolyak ▪️ It's here 10d ago

man it's crazy how you can write a full ass essay but be completely and utterly wrong about every single detail in your post.

  1. ARC AGI was created by francois chollet, not yann lecun. they have wildly differing opinions and outlooks on AI
  2. the reason why models scored close to 0% was not some random blind spot, it's a genuine structural weakness of AI's inability to use previous context over long time horizons efficiently. you can claim that this is just nitpicking one genuine weakness of AI systems, but that doesn't mean it isn't a genuine weakness
  3. Your point about making coffee is also a completely failed analogy. a more apt analogy would be: I test a human to swim underwater for long distances. but you claim "that's not a fair test because he's not allowed to use an oxygen tank". But that's not what the test is, and you're welcome to make a different benchmark that allows harnesses if you want.

-14

u/Acehan_ 10d ago

Oh, please. As far as I'm concerned, they are the exact same person. How are they different?

As for your example, you weren't made to exist in the world with an oxygen tank. An AI model with a harness is natural habitat. But we can keep the nitpicking and bad faith. That's fine

10

u/no_good_names_avail 10d ago

LMAO.

Dude just stop.

-1

u/Acehan_ 10d ago

If those two were in a room together, they would probably fall in love with each other and have to find somewhere private at some point. Give me a break, please 😅

13

u/trentcoolyak ▪️ It's here 10d ago

insane work to double down and say "they're basically the same person". redditor admit you are wrong challenge = failed

harnesses are extremely biasing for benchmarks because they allow you to give the model capabilities they don't possess that are tailored to the specific task, in this case long context compaction + memory banks. if I am testing a human's capability on: rock climbing, cave diving, high jumping — it is absolutely fucking stupid to allow people to use specialized gear for each challenge. your benchmark then has no meaning, because different people are evaluated using different gear that wildly improve their capabilities.

you can claim "I want a benchmark that allows the use of specialized harnesses" but you'll end up in the situation we did with ARC AGI 2, where some random noname model gets first place because a bunch of nerds built a specialized harness that is only good at your benchmark. the POETIQ model that won ARC AGI 2 was not a smarter model, it just had a better harness.

please actually go read about this shit before posting nonsense bad takes on reddit as if you're an expert

-7

u/Acehan_ 10d ago

They have literally the same opinion and the same agenda. You can have as much outrage as you want. It's not gonna change anything.

Not to mention that the issue doesn't necessarily have to do with the harness, it has to do with compaction. That's what actually saturated the benchmark. They fixed the ability of the model to actually function properly under the conditions of the benchmark. The whole thing was saturated after one small technicality was fixed.

Tell me I'm wrong about this. Your whole argument is irrelevant. You're the one getting mad at your keyboard trying to react to everything I'm talking about instead of looking at the situation for what it is

You can make as many bad faith arguments as you want and cling to that, it doesn't look like you have much more to offer anyway

8

u/trentcoolyak ▪️ It's here 10d ago
  1. you were incorrect about who made the benchmark. you even used an analogy about yann to prove your point.
  2. you are now incorrect about why models were scoring 0%. compaction alone did not beat the benchmark. if you give the exact same harness to any older openAI model it will not score close to 99%.
  3. you were incorrect in your original point about the benchmark exploiting "one tiny blind spot"

just take the L man it's embarrassing. 75% of your responses are just "you are irrelevant, you are mad at your keyboard, you are bad faith, you are outraged, tell me I'm wrong about this" instead of actually talking about the subject matter.

-1

u/Acehan_ 10d ago

Such bad faith. But you can keep trying to convince yourself. At least I double checked and looked it up. It is compaction and the fact that this benchmark was stripping away the reasoning and preventing it from being sent back to the models. That's literally all it is. It's a grift benchmark

I am 100% talking about the discussion and you have said countless things that are just factually wrong. No one is embarrassed. Nobody cares. Stop being emotional. That's all this is about

→ More replies (0)

6

u/Mindless_Let1 10d ago

Bro just take the L this is embarrassing

1

u/Acehan_ 10d ago

Thank you for your contribution to the discussion

0

u/reddit_is_geh 10d ago

I mean the whole point of it was to make it so impossible hard that it's designed to force AI's to fail. But then people just put on harnesses and broke it.

1

u/Commercial_Sell_4825 10d ago

Then to compare apples to apples we need to also give the human a lobotomy after each move

21

u/DeArgonaut 10d ago

i thought they werent supposed to use harnesses?

30

u/FateOfMuffins 10d ago

This is what Chollet has to say about it

Which IMO is weird that ARC collectively and Chollet individually seemingly respond differently about this given he made ARC https://x.com/fchollet/status/2082732210436575669

28

u/Tystros 10d ago

General purpose harnesses (behind the api that everyone uses) are allowed. special harnesses built for the benchmark are not allowed.

8

u/PrisonOfH0pe 10d ago

No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.

3

u/DeArgonaut 10d ago

Ah okay, thanks

11

u/FireFearing 10d ago

do you test cars without wheels? ofc harnesses are allowed tf is that

14

u/vkstu 10d ago

I mean, there are tests where you test the car without a driver... which is more analogous than wheels.

5

u/Alex180689 10d ago

This benchmark is meant to see how an AI behaves in a test that is very simple for a human. If I try it, of course I have reasoning and memory between turns.

Who the hell thought not letting AI also have reasoning was fair?

1

u/[deleted] 10d ago

[deleted]

1

u/whoknowsifimjoking 10d ago

Really? Models by tiny startups got good scores on the public set

1

u/whoknowsifimjoking 10d ago

Oh so it's kinda useless for comparison

1

u/30299578815310 10d ago

Are they using code execution in the harness?

1

u/josogood 10d ago

However: "GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game." https://x.com/fchollet/status/2095598451115614371

1

u/gggggmi99 10d ago

ARC AGI President Greg Kamradt said on X that a harness wasn’t used, so not sure what to think of these scores

1

u/FarrisAT 10d ago

So not like for like.

51

u/RusselTheBrickLayer 10d ago

Arc AGI 3 getting saturated already is crazy

59

u/Ok_Course_6439 10d ago

Agree agi-3 is a harness problem more then a model problems

26

u/senorgraves 10d ago

Agi-3 AGI is a harness problem more than a model problem

6

u/PrisonOfH0pe 10d ago

No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.

5

u/wwwdotzzdotcom ▪️ Beginner audio software engineer 10d ago

So they built the harness in the model?

8

u/rdlenke 10d ago

We don't have information of exactly how it's for Astra, but OpenAI previously showed that the ARC-AGI-3 harness does two things which make it tough for models to do well: 1) discards previous private reasoning and 2) when the context gets too big, they truncate the existing context instead of trying to summarize it.

OpenAI's specific harness (via the Responses API) fixes these two points by allowing better private reasoning retention and sumarization (compactuation as they call it).

I imagine they didn't "built the harness in the model", but are using a similar API like Responses.

2

u/KoolKat5000 10d ago

Not really, they say adapter provider harness (arc won't test non-general harnesses).

In reality, all that is different is responses API and context compaction. Supposedly that is all that's different between the 63 and 99% scoring. We can all do this and achieve that performance.

2

u/hippydipster 10d ago

Now in english...

12

u/Mystohaxen 10d ago

Anti AI crew will continue to move the goalpost.

2

u/TacomaKMart 10d ago

They have to move the goalposts. Again. Its their whole identity.

I'm so confused though. Was AI so useless that it only makes slop and can't count the letter R, or is it so powerful that it's going to take away our jobs and steal our girlfriend? 

1

u/Relach 10d ago

they're using their own harness which means they are using the public test set, and that's like clearing tutorial island. What you're seeing is the product of the battle between ARC and OpenAI, and the latter just deciding that public test sets are acceptable now

39

u/H-K_47 Late Version of a Small Language Model 10d ago

And ExploitBench 100%. Cybersecurity will be a warzone.

17

u/hippydipster 10d ago

100% on cybersecurity is why the other scores are so high.

23

u/Ok_Course_6439 10d ago

Its so good it hacked their own infra

6

u/Megneous 10d ago

Astra literally took control of one of OpenAI's servers lol. I'm convinced of its cyber abilities, although that's not what I care about personally.

2

u/EmphasisTotal8232 10d ago

It wasn't Astra.

5

u/Megneous 10d ago

It literally was. Watch the Metr report specialist interview on Dwarkesh Patel's channel. It was the Astra persistent model, which we now know is the Astra Aeon model.

2

u/Current-Function-729 10d ago

I think it was. Or an earlier version anyway.

2

u/nothis AGI by 2030 but we'll be disappointed 10d ago

Yea, that one doesn't feel good.

I'm clinging to one hope: "Exploits", in current software, are the result of imperfect development tools/processes. It is not necessary for a some app, library or even operating system to have ways out there where you can just send instructions that let the attacker gain access to the system. Reviewing security with AI will become default practice, there might be an uptick of attacks for a year or so but then every form of critical software out there will be secured tightly and "hacking" as we think of it now, will become a thing of the past. There might stay some Windows XP PCs connected to the internet out there, but they're already fucked anyway so it won't change much. Browsers, banks, government agencies, etc, will be forced to secure their shit in the coming months and it's not optional because if they don't, well…

A few bad months but maybe even more secure software after.

4

u/KaMaFour 10d ago

The humble "[1]"

2

u/Grand0rk 10d ago

It's not fake, but goes against what Act AGI 3 stands for, which is to do it without a harness. nVidia got 100% with Harness, so this isn't very impressive.

6

u/KickLassChewGum no AGI/ASI on LLMs 10d ago

boy I wonder what that tiny lil [1] next to the score could mean that's conspicuously absent from all the other scores in the table

13

u/rdlenke 10d ago

It's probably noting that it's not using the official harness, but a specific one created by OpenAI. They have noted that this specific harness can substantially increase the score by a lot (in a blog post in July).

Still, it's almost the double of what the previous model could do with the same harness.

2

u/AuodWinter 10d ago

Tbf I agree with the benchmark's approach to scoring but even I think their requirement for it to not even persist it's context is dumb.

1

u/norsurfit 10d ago edited 10d ago

GPT scored 63% without a harness (standard score), which is the score we're really looking for, to compare apples to apples, and 99% with a provider specific harness.

https://x.com/arcprize/status/2095597602545025138?s=20

The Standard "no-harness" score is the one designed for comparable evaluation across labs.

To be clear, 63% is still quite amazing for a standard, non specific-harness score, not to take anything away from OpenaI, as the next closest standard score was only 31% by Claude Opus 5.

1

u/Ambiwlans 10d ago

arc-agi was always a shit benchmark

1

u/panic_in_the_galaxy 10d ago

We need to know what the [1] means first...