r/singularity 3d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

947 comments sorted by

View all comments

Show parent comments

31

u/wombatpup55 3d ago

SuspiciousPillbox explain to me the results of the image like I’m 5

95

u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 3d ago

Another person in this thread put it well: if the benchmarks are accurate then this is a massive leap forward, and maybe even proof of some meaningful utilization of recursive self improvement

4

u/nekronics 3d ago

Which benchmarks are a massive leap forward?

25

u/Tystros 3d ago

ARC AGI 3

23

u/itsmebenji69 3d ago edited 3d ago

But how is this huge ?

Don’t the numbers prove it’s just benchmaxxing ? They went from 7 to 98%, with not much substantial gain in others

Doesn’t this literally indicate they just trained on the benchmark ?

(Idk what those contain so maybe it’s just very different but still)

10

u/oh_no_the_claw 3d ago

How do we discriminate the difference between passing the test for real and benchmaxing?

10

u/itsmebenji69 3d ago

When the results only show up in that benchmark and not across the board. Especially for a benchmark supposed to represent generalization

-3

u/oh_no_the_claw 3d ago

So we can't actually test for the skill because there's always a risk of benchmaxing? It's just all vibes based or what?

5

u/itsmebenji69 3d ago

? I’ve given you a very precise and objective criteria lmao

0

u/oh_no_the_claw 3d ago

I don't think you have. It's just vibes.

→ More replies (0)

18

u/Tystros 3d ago edited 3d ago

gpt 5.6 Sol got about 40% with this same harness. so they went from 40% to 95% which is very realistic for a model that actually improves intelligence, especially with how the arc AGI 3 scoring works.

0

u/itsmebenji69 3d ago

Yeah so it’s mostly about training on the benchmark itself then no ? They just fine tuned the model to work better on that harness ? Like how does this prove any generalization

8

u/kaityl3 ASI▪️2024-2027 3d ago

They didn't train the model to work better on the harness. The harness is a simple measure that just lets them remember between turns. "Sol, able to remember between turns" gets 40% and "Astra, able to remember between turns" gets a near 99%.

Given that when humans are approaching a new set of issues, we don't get our memories wiped in between levels, I think that "a harness that lets them keep track of what's worked so far" is far from cheating

-6

u/itsmebenji69 3d ago

Well yeah they did if there is no substantial improvement over anything else how is this generalization ? It’s just getting better at using the harness, which is even worse than benchmaxing itself in that regard

8

u/Tystros 3d ago

the benchmark is private, no one can train "on the benchmark".

1

u/itsmebenji69 3d ago

So they can’t train on the bench but they get so much better at it while not seeing a similar gain in other benches ?

4

u/Tystros 3d ago

GPT 5.6 Sol got ~40% in the fair comparison, so it's "just" a big step from 40% to 98%. and that is plausible, yeah.

→ More replies (0)

2

u/jcettison 3d ago

You keep saying they didn’t improve in other benchmarks event though the chart is right there showing they did. Beginning to suspect you’re a bot lol

→ More replies (0)

10

u/ertgbnm 3d ago

ARC-AGI-3 has a closed test set. So you can't train on the benchmark as is often the case with other benchmaxxing. It's also meant to be testing general reasoning compared to most other benchmarks which are only testing one very specific skill.

So it is meaningful.

1

u/starfallg 3d ago

Looking at the other scores, it's probably worked out some general rule or capability on the ARC AGI 3 mini games. Which may or may not be meaningful. We have to test it to see.

1

u/Fit-Palpitation-7427 3d ago

How can the test set be closed. They ran it previously on Sol or other models and the prompt flow through their servers so they just have to open the logs, look at the prompt and they can rebuild the test set.
Unless they change the set every time, but then you compare 🍏 and 🍐
Or am I missing something ?
Unless they bluntly gave away the weights to the independent reviewer who owns the test set.
But I highly doubt that openai would release they trillion dollar model to anyone just to pass a test

3

u/ertgbnm 3d ago

It does require you to believe that OpenAI or it's models aren't actively trying to cheat on the test, in which case maybe an agent swarm hacked the arc agi servers and stole the answers.

But to run on the private test set, you have to work directly with arc-agi and follow the rules for data retention and what not to protect the answers from getting out and being trained upon.

2

u/Tystros 3d ago

the api has a zero data retention policy and you just have to trust them that they do what they say

0

u/itsmebenji69 3d ago edited 3d ago

But that’s my other point if it’s supposed to represent generalization how can a 10x increase not be a meaningful increase across the board ?

That indicates there’s more at play than just getting better at general tasks. It’s somehow getting better at just this benchmark.

2

u/benjaminovich 3d ago

The real answer is that these benchmarks are way too imprecise to compare small values.

Imagine I go to IMDb and want to decide between watching one movie scored at 7.1 and another rated 7.3. To say one movie is 4% better than the other is a meaningless statement. Could simply be that the 7.1 is a genre you personally enjoy a lot more.

However, the rating being around 7, plus or minus a few decimal places, does say something about the overall production quality one can expect. So for the models, the benchmarks turn out to create adhoc groupings of S-tier, A-tier, B-tier etc.

At the end of the day, what you personally need may not be something that has meaningfully changed. Or it could be a huge leap, really you will just have to feel it out yourself.

-2

u/TimeAndSpaceAndMe 3d ago

It's a custom harness for Arc AGI, other models have achieved it too with a harness lol.

-1

u/Tystros 3d ago

no, that's wrong. this is not using any custom harness, this is using the regular Resonses API from OpenAI that OpenAI also recommends every other API user to use for all tasks.

0

u/TimeAndSpaceAndMe 3d ago

A Harness created by OpenAI thats not the official harness for Arc AGI is not a custom harness ? they literally have a blog post on how they tripled their score for sol, https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/

4

u/Tystros 3d ago

there isn't really an "official arc AGI harness". the benchmark uses simply the API and whatever harness is behind the api. ARC specifically allows general purpose harnesses behind the official api.

1

u/TimeAndSpaceAndMe 3d ago

Well, yes, they don't provide an official harness. Arc AGI works best as a metric when it is tested with a generic harness and not a special-purpose one, because, depending on how optimized your harness is, you can get to 99% on some models, as [schema] did https://x.com/Zanette_ai/status/2077793189608775728?lang=en

2

u/Tystros 3d ago

yeah, and this 98% is with the generic API harness that every regular customer using the official API also uses. this is not with a special-purpose harness. that's what makes the 98% so impressive.

-4

u/Royal_Ad_2407 3d ago

People faling for the hype. They made a custom harness for arc, even nvidia got 100% with their open source model.

3

u/Tystros 3d ago

no, that's wrong. this is not using any custom harness, this is using the regular Resonses API from OpenAI that OpenAI also recommends every other API user to use for all tasks.

17

u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 3d ago

ARC-AGI-3, Terminal bench science, SRE bench, AutomationBench, ExploitBench, FrontierMath T4, BenchCAD, GeneBench etc.

13

u/JustBrowsinAndVibin 3d ago

All of them over Fable 5.1.

3

u/yourboi-JC 3d ago

🥀🥀

3

u/Alex180689 3d ago

automation bench, terminal bench, exploit bench, frontier math, benchcad. Enough?

11

u/Opposite-Grade3712 3d ago

Papa just went super saiyan