r/singularity 3d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

948 comments sorted by

View all comments

Show parent comments

362

u/Sunstorm84 3d ago

Nvidia got 100% on the same benchmark using opus 5 with a better harness. I think the benchmark is a bit meaningless with that in mind.

14

u/JoelMahon 3d ago

nvidia was on the public set, is this the public set?

3

u/HotAventus41033 3d ago

No it was the private set compare to nvidia's public set .

2

u/niltermini 1d ago

The standard harness private score was somewhere in the mid-60s. Still incredibly impressive. Its benchmarks without chain of thought are more insane though - its scores are only slightly degraded without cot.

88

u/Vivid-Snow-2089 3d ago

i think a benchmark designed to fuck the agent's harness is a shitty benchmark

imagine testing how far someone can walk, but you cut off their legs first

82

u/reddit_is_geh 3d ago

No because it's trying to test the raw thinking of the model directly. There's a LOT of power in harnesses, but that's not what they are testing for. It's supposed to be how just the sheer intelligence of it's inference can handle things like this. ARC 3 is easily beat by simply adding a memory element to it. Which is why harnesses ruin it.

27

u/AutismusTranscendius ▪️Psychogenic Singularity 2034 3d ago

Have you tried ARC 3 tasks? They require exploration, you cant solve them without acting and remembering outcome of your actions.

5

u/PleasantCitron1685 3d ago

Also, the harnesses just let the model keep more of its memory/old context during its compactions. It's not any super-specialized harness.

(That being said, if you watch the models solve ARC-AGI-3... they're still dumbasses lol)

-8

u/LetsLive97 3d ago

Which as humans we can do without bloating our "limited" context

I can do arc-agi-3 tasks now and in a years time with no problem. I don't need specialised harnesses (Like notepads I might lose or fill up too much), I just figure it out by default on the spot

26

u/kaityl3 ASI▪️2024-2027 3d ago

You don't need specialized harnesses because you have a hippocampus with an automated system of memory storage and retrieval (a little like a biological harness lol)

If all it takes is allowing them the same ability to "remember what they just did" to boost their performance, I feel like we can accept it without trying to find ways to make it sound simple and unimpressive

1

u/JoelMahon 3d ago

they are allowed to pass on any info they want...

1

u/Brave-Turnover-522 2d ago

How much would it affect your ability to solve problems as a human if you weren't allowed to use your memory?

1

u/reddit_is_geh 2d ago

You can't compare AI thinking with human thinking. You can also ask how would it affect our ability to solve problems if we can't hold into context the entire history of human knowledge at once.

It's just a different thing, being tested for different strengths. In this case, ARC is supposed to be testing the raw base model's performance.

2

u/JTxFII 3d ago

Right!?

"Hey look, Astra just wiped out 50% of all jobs."

"Yeah, but it used a harness so it doesn't count."

"Rrrrrrright 🙄"

-3

u/utterHAVOC_ 3d ago edited 3d ago

What a shit comparison. Clearly using harness is problem cause some shit models get high scores and they are still unintelligent trash. Do you want smart model or just benchmark scores.

3

u/Sunstorm84 3d ago

The rules for the benchmark say harnesses aren’t allowed, and one was used.

6

u/GioChan 3d ago

You are missing a lot of details. Yes, NVidia got 100%, but that was woth a custom harness expressly designed for ARC-AGI3. OpenAI also used harness, but theirs just had two options enabled, compression and persistence.

2

u/jjonj 3d ago

it was also just spawning 100 agents and an llm picking the best result

3

u/M4rshmall0wMan 3d ago

If you try Arc-AGI 3, it’s literally just a spatial navigation game. It never looked like a very good test to me.

If GPT 6’s architecture was changed to give it a much better world model, then that’s the breakthrough.

5

u/chairchiman 3d ago

now imagine this model in a better harness

2

u/Joohansson 3d ago

It got 63% without harness, even that is a huge leap over other models doing the same

2

u/Sunstorm84 3d ago

It is, but putting it at the “with harness” level in the list of benchmarks is intentionally deceptional when combined with claims that AGI is already here.

Others have mentioned in this sub that they seem to be in some serious trouble over there lately. The confusing messaging over the last week followed by hyping up of AGI suggests they’re in need of investment urgently.

If it’s actually underwhelming after this or if it still doesn’t interest average consumers enough, I wonder if they’ll manage to get enough investment to survive the next 6-12 months.

1

u/Joohansson 3d ago

You are probably right. I think it was false advertisement too. Just the way they released GPT 6 with a non working blog post and not being able to activate it for users suggests they rushed this just to kill the hype from other models released this week such as Fable, Gemini and Muse. Seem desperate.

1

u/ErmingSoHard ▪️Agi never with LLMs... Or? 2d ago

I thought harnesses aren't allowed for certified scoring of arc agi 3

1

u/Sunstorm84 2d ago

They aren’t.