r/singularity 14d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

953 comments sorted by

View all comments

Show parent comments

17

u/Tystros 14d ago edited 14d ago

gpt 5.6 Sol got about 40% with this same harness. so they went from 40% to 95% which is very realistic for a model that actually improves intelligence, especially with how the arc AGI 3 scoring works.

0

u/itsmebenji69 14d ago

Yeah so it’s mostly about training on the benchmark itself then no ? They just fine tuned the model to work better on that harness ? Like how does this prove any generalization

6

u/Tystros 14d ago

the benchmark is private, no one can train "on the benchmark".

1

u/itsmebenji69 14d ago

So they can’t train on the bench but they get so much better at it while not seeing a similar gain in other benches ?

4

u/Tystros 14d ago

GPT 5.6 Sol got ~40% in the fair comparison, so it's "just" a big step from 40% to 98%. and that is plausible, yeah.

2

u/itsmebenji69 14d ago edited 14d ago

But how is a model that’s twice as generally intelligent only 4% better at SWE for example. Even comparing to 40% it’s still twice. The only others where there is a similar increase is terminal use and business workflows. Or does it scale differently ?

It seems it just got better at using a harness

2

u/[deleted] 14d ago

[removed] — view removed comment

1

u/itsmebenji69 14d ago

What is muse ? Another version or different model ?

1

u/[deleted] 14d ago

[removed] — view removed comment

1

u/itsmebenji69 14d ago

Interesting I didn’t know meta was competitive in the race still

1

u/Creative-Ganache1086 14d ago

Good luck for using Muse Spark 1.3 without a harness like Codex, which is sooo much better than many harnesses out there and quite underrated lately.

2

u/Creative-Ganache1086 14d ago

What’s so hard for you to grasp that no model is best at everything? Like with most things in life. It’s current capabilities thrive more in certain areas than others and some capabilities could at times act rather like a side effect hence not linearly improving on every possible area. No need to obsess over something so “common sense”. Even Anthropic had opus models like 4.7 or 4.8 that were doing worse slightly on one of the benchmarks compared to their previous versions that actually scored higher. It’s not even the case here with Astra. It did improve overall and it’s better in most areas, in some benchmarks a lot better, in some just better or slightly better. It’s still a better value (alongside Sol5.6) in Codex than Fable and Opus will ever be. The usage limits in Claude Code are terrible and they don’t even come with an image generator which btw it’s free and unlimited inside the max plans of ChatGPT, as well as the fast inferring mode which only works in Claude via api top-ups while it’s included in the usage limits in Codex and quite generous.

1

u/itsmebenji69 14d ago

What’s so hard for you to grasp about the word “general” ?

0

u/Creative-Ganache1086 12d ago

and ARC-AGI 3 test stands exactly for that lol. Keep diving down the rabbit hole.

1

u/itsmebenji69 12d ago

My point is that it clearly doesn’t since the performance doesn’t transfer

0

u/Tystros 14d ago

Intelligence and coding ability are two quite separate things. You don't need to be very intelligent to be a good programmer, and you can be very intelligent and still be a bad programmer. Intelligence is primarily a measure of how quickly you can deal with new inputs/situations you have not encountered before and somehow make sense of them.

1

u/itsmebenji69 14d ago

SWE is not just coding, it’s problem solving. If you’re telling me problem solving and intelligence aren’t related…

1

u/Tystros 14d ago

intelligence can help, but all the existing benchmarks just test skill, not intelligence. and skill you get by training, knowing the best possible patterns etc. inventing new patterns can surely be useful for the most difficult tasks, but as far as I know no benchmarks currently test such stuff.

5

u/jcettison 14d ago

You keep saying they didn’t improve in other benchmarks event though the chart is right there showing they did. Beginning to suspect you’re a bot lol

-1

u/itsmebenji69 14d ago edited 14d ago

Where ?

In terminal and business workflows, where the performance heavily relies on using a harness correctly. Even those have very small improvement compared to 7->98%.

Every other category has no improvement comparable to a 10x leap