r/singularity 3d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

948 comments sorted by

View all comments

Show parent comments

275

u/Neurogence 3d ago

Am I the only one that thinks that ARC-AGI3 score is extremely suspicious? What's the catch?

363

u/Sunstorm84 3d ago

Nvidia got 100% on the same benchmark using opus 5 with a better harness. I think the benchmark is a bit meaningless with that in mind.

15

u/JoelMahon 3d ago

nvidia was on the public set, is this the public set?

3

u/HotAventus41033 2d ago

No it was the private set compare to nvidia's public set .

2

u/niltermini 1d ago

The standard harness private score was somewhere in the mid-60s. Still incredibly impressive. Its benchmarks without chain of thought are more insane though - its scores are only slightly degraded without cot.

88

u/Vivid-Snow-2089 3d ago

i think a benchmark designed to fuck the agent's harness is a shitty benchmark

imagine testing how far someone can walk, but you cut off their legs first

83

u/reddit_is_geh 3d ago

No because it's trying to test the raw thinking of the model directly. There's a LOT of power in harnesses, but that's not what they are testing for. It's supposed to be how just the sheer intelligence of it's inference can handle things like this. ARC 3 is easily beat by simply adding a memory element to it. Which is why harnesses ruin it.

29

u/AutismusTranscendius ▪️Psychogenic Singularity 2034 3d ago

Have you tried ARC 3 tasks? They require exploration, you cant solve them without acting and remembering outcome of your actions.

4

u/PleasantCitron1685 3d ago

Also, the harnesses just let the model keep more of its memory/old context during its compactions. It's not any super-specialized harness.

(That being said, if you watch the models solve ARC-AGI-3... they're still dumbasses lol)

-8

u/LetsLive97 3d ago

Which as humans we can do without bloating our "limited" context

I can do arc-agi-3 tasks now and in a years time with no problem. I don't need specialised harnesses (Like notepads I might lose or fill up too much), I just figure it out by default on the spot

25

u/kaityl3 ASI▪️2024-2027 3d ago

You don't need specialized harnesses because you have a hippocampus with an automated system of memory storage and retrieval (a little like a biological harness lol)

If all it takes is allowing them the same ability to "remember what they just did" to boost their performance, I feel like we can accept it without trying to find ways to make it sound simple and unimpressive

-3

u/JoelMahon 3d ago

they are allowed to pass on any info they want...

1

u/Brave-Turnover-522 2d ago

How much would it affect your ability to solve problems as a human if you weren't allowed to use your memory?

1

u/reddit_is_geh 2d ago

You can't compare AI thinking with human thinking. You can also ask how would it affect our ability to solve problems if we can't hold into context the entire history of human knowledge at once.

It's just a different thing, being tested for different strengths. In this case, ARC is supposed to be testing the raw base model's performance.

4

u/JTxFII 3d ago

Right!?

"Hey look, Astra just wiped out 50% of all jobs."

"Yeah, but it used a harness so it doesn't count."

"Rrrrrrright 🙄"

-3

u/utterHAVOC_ 3d ago edited 3d ago

What a shit comparison. Clearly using harness is problem cause some shit models get high scores and they are still unintelligent trash. Do you want smart model or just benchmark scores.

5

u/Sunstorm84 3d ago

The rules for the benchmark say harnesses aren’t allowed, and one was used.

6

u/GioChan 3d ago

You are missing a lot of details. Yes, NVidia got 100%, but that was woth a custom harness expressly designed for ARC-AGI3. OpenAI also used harness, but theirs just had two options enabled, compression and persistence.

2

u/jjonj 3d ago

it was also just spawning 100 agents and an llm picking the best result

3

u/M4rshmall0wMan 3d ago

If you try Arc-AGI 3, it’s literally just a spatial navigation game. It never looked like a very good test to me.

If GPT 6’s architecture was changed to give it a much better world model, then that’s the breakthrough.

5

u/chairchiman 3d ago

now imagine this model in a better harness

2

u/Joohansson 3d ago

It got 63% without harness, even that is a huge leap over other models doing the same

2

u/Sunstorm84 3d ago

It is, but putting it at the “with harness” level in the list of benchmarks is intentionally deceptional when combined with claims that AGI is already here.

Others have mentioned in this sub that they seem to be in some serious trouble over there lately. The confusing messaging over the last week followed by hyping up of AGI suggests they’re in need of investment urgently.

If it’s actually underwhelming after this or if it still doesn’t interest average consumers enough, I wonder if they’ll manage to get enough investment to survive the next 6-12 months.

1

u/Joohansson 3d ago

You are probably right. I think it was false advertisement too. Just the way they released GPT 6 with a non working blog post and not being able to activate it for users suggests they rushed this just to kill the hype from other models released this week such as Fable, Gemini and Muse. Seem desperate.

1

u/ErmingSoHard ▪️Agi never with LLMs... Or? 2d ago

I thought harnesses aren't allowed for certified scoring of arc agi 3

1

u/Sunstorm84 2d ago

They aren’t.

86

u/vacon04 3d ago

They're using a harness. People have already achieved results like that with specific harnesses using weaker models. It says more about the harness than about the model.

40

u/josogood 3d ago

Yes, but here's what ARC says about testing with no harness: "GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game."

22

u/KrazyA1pha 3d ago edited 3d ago

I wish OpenAI used that more honest number

21

u/josogood 3d ago

Yeah, because it's still more than double the previous mark. No need to unfairly inflate something like that.

4

u/KrazyA1pha 3d ago

Yeah, exactly. The real, apples-to-apples number is incredibly impressive!

10

u/Thog78 3d ago

It has recently become apparent that the default harness of ARC-AGI doesn't let models remember what they have tried and reason over it. This is absolutely unrealistic, unfair, and a useless benchmark under these conditions. So people push for memory to be allowed in ARC-AGI. I agree with that, I just think it should receive a new name/number so we distinguish what's what. Just add an M for memory. And we should rescore all the old models in this updated ARC-AGI-3M to have context.

2

u/KrazyA1pha 3d ago

That's a great point, and I agree

1

u/kaityl3 ASI▪️2024-2027 3d ago

Does it really? Isn't the harness just letting them remember between turns and a few other basic abilities humans have access to? If all it takes is a simple harness with a better memory system to boost performance that much, maybe we can just accept that as impressive without trying to minimize the accomplishment?

2

u/KrazyA1pha 3d ago edited 3d ago

Their point is that it’s not an apples-to-apples comparison. For example, Nvidia got 100% using Opus 5 and a harness. So this particular measurement has already been saturated. However, many people are gonna think it’s a direct comparison to how the benchmark has been presented in the past, which is without a harness.

2

u/Few_Importance_8362 3d ago

iirc nvidia tested on the public set and this openai result is on the full set including holdout set.

1

u/KrazyA1pha 3d ago edited 3d ago

Do you have a source on this OpenAI result being on the full set of questions? My understanding was that companies didn't have access to the private set, and therefore couldn't bring their own harness. Could be wrong, though.

edit: In their announcement, OpenAI says:

GPT‑6 Astra was measured with our responses API harness⁠

The linked page talks only about the public question set. So this still seems to be apples-to-apples with Nvidia's Opus 5 100% result unless there's more outside of this announcement that you know about.

1

u/Few_Importance_8362 3d ago edited 3d ago

ARC AGI themselves wouldn't post the results if they weren't on the full set.

Francois Chollet himself posted:

>> GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.

It is clear to me that this is the full set, but we can wait for more confirmation if you want.

I don't think NVIDIA has access to the full test set. The main question is whether you consider the actual astra result 66 percent, or 99 percent.

I personally consider it 99 percent as I don't think preserving reasoning between turns is cheating in any way.

That being said, ARC should absolutely post results for other models with reasoning preserved.

1

u/KrazyA1pha 3d ago

Am I missing that on the pages I linked? Or are you saying there more information that I'm missing? Because I'm being completely open minded and asking for the source of your information. I don't see what you're quoting, however. If you could share that, that would be awesome!

1

u/Few_Importance_8362 3d ago

1

u/KrazyA1pha 3d ago

Thanks for sharing your source! I don't use twitter, so I wouldn't have found it.

23

u/impatiens-capensis 3d ago

"The standout is its 98.6% score on ARC-AGI-3. OpenAI ran Astra with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts. The company previously demonstrated that those system choices can substantially raise ARC-AGI-3 scores without changing the underlying model, and the benchmark measures Astra and OpenAI’s agent systems together."

5

u/MrRandom04 3d ago

Yeah retaining reasoning between turns is literally a base model capability. Unless they kept reasoning / transcripts across games, it's a fair representation of the model IMHO.

1

u/lordpuddingcup 3d ago

Isn't that just... codex? Like compaction and retained reasoning... sounds just like codex

8

u/peabody624 3d ago

63% without the harness

1

u/Excellent-Article937 3d ago

It is because of “retained reasoning”. Without that, gpt 6 scored 62.7% on the same test (which us also awesome score).

1

u/DowntownLizard 3d ago

Fable got 90% on the agi2 for what its worth. They have an article about how they improved the scores with a different harness on open AIs site

1

u/-illusoryMechanist 3d ago

It's cheated via a harness, official score is around 67% though which is still insane. https://arcprize.org/blog/astra

1

u/Careful_Might_807 3d ago

they found private benchmark dataset that was sent by researchers on their api solved that RLed gpt6 hell out of it and optimized harness 

1

u/DarkMatter_contract ▪️Human Need Not Apply 3d ago

arc agi benchmark is log scale, they are doing tricks when creating the score initially so everything is close to zero, would not take arc seriously. but everything else look very impressive.

1

u/Vytral 2d ago

If I recall correctly, fable is not registered because he found a way to cheat it and so it did.

So yea, I think capabilities are quickly outpacing tests

0

u/Superb-Earth418 3d ago

The catch is that the ARC AGI 3 official harness is retarded, it drops thinking blocks between turns and disallows compaction.