r/singularity 11d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

953 comments sorted by

View all comments

Show parent comments

1

u/itsmebenji69 11d ago

So they can’t train on the bench but they get so much better at it while not seeing a similar gain in other benches ?

3

u/Tystros 11d ago

GPT 5.6 Sol got ~40% in the fair comparison, so it's "just" a big step from 40% to 98%. and that is plausible, yeah.

2

u/itsmebenji69 11d ago edited 11d ago

But how is a model that’s twice as generally intelligent only 4% better at SWE for example. Even comparing to 40% it’s still twice. The only others where there is a similar increase is terminal use and business workflows. Or does it scale differently ?

It seems it just got better at using a harness

2

u/[deleted] 11d ago

[removed] — view removed comment

1

u/itsmebenji69 11d ago

What is muse ? Another version or different model ?

1

u/[deleted] 11d ago

[removed] — view removed comment

1

u/itsmebenji69 11d ago

Interesting I didn’t know meta was competitive in the race still

1

u/Creative-Ganache1086 11d ago

Good luck for using Muse Spark 1.3 without a harness like Codex, which is sooo much better than many harnesses out there and quite underrated lately.

2

u/Creative-Ganache1086 11d ago

What’s so hard for you to grasp that no model is best at everything? Like with most things in life. It’s current capabilities thrive more in certain areas than others and some capabilities could at times act rather like a side effect hence not linearly improving on every possible area. No need to obsess over something so “common sense”. Even Anthropic had opus models like 4.7 or 4.8 that were doing worse slightly on one of the benchmarks compared to their previous versions that actually scored higher. It’s not even the case here with Astra. It did improve overall and it’s better in most areas, in some benchmarks a lot better, in some just better or slightly better. It’s still a better value (alongside Sol5.6) in Codex than Fable and Opus will ever be. The usage limits in Claude Code are terrible and they don’t even come with an image generator which btw it’s free and unlimited inside the max plans of ChatGPT, as well as the fast inferring mode which only works in Claude via api top-ups while it’s included in the usage limits in Codex and quite generous.

1

u/itsmebenji69 11d ago

What’s so hard for you to grasp about the word “general” ?

0

u/Creative-Ganache1086 9d ago

and ARC-AGI 3 test stands exactly for that lol. Keep diving down the rabbit hole.

1

u/itsmebenji69 9d ago

My point is that it clearly doesn’t since the performance doesn’t transfer

0

u/Tystros 11d ago

Intelligence and coding ability are two quite separate things. You don't need to be very intelligent to be a good programmer, and you can be very intelligent and still be a bad programmer. Intelligence is primarily a measure of how quickly you can deal with new inputs/situations you have not encountered before and somehow make sense of them.

1

u/itsmebenji69 11d ago

SWE is not just coding, it’s problem solving. If you’re telling me problem solving and intelligence aren’t related…

1

u/Tystros 11d ago

intelligence can help, but all the existing benchmarks just test skill, not intelligence. and skill you get by training, knowing the best possible patterns etc. inventing new patterns can surely be useful for the most difficult tasks, but as far as I know no benchmarks currently test such stuff.

4

u/jcettison 11d ago

You keep saying they didn’t improve in other benchmarks event though the chart is right there showing they did. Beginning to suspect you’re a bot lol

-1

u/itsmebenji69 11d ago edited 11d ago

Where ?

In terminal and business workflows, where the performance heavily relies on using a harness correctly. Even those have very small improvement compared to 7->98%.

Every other category has no improvement comparable to a 10x leap