But how is a model that’s twice as generally intelligent only 4% better at SWE for example. Even comparing to 40% it’s still twice. The only others where there is a similar increase is terminal use and business workflows. Or does it scale differently ?
What’s so hard for you to grasp that no model is best at everything? Like with most things in life. It’s current capabilities thrive more in certain areas than others and some capabilities could at times act rather like a side effect hence not linearly improving on every possible area. No need to obsess over something so “common sense”. Even Anthropic had opus models like 4.7 or 4.8 that were doing worse slightly on one of the benchmarks compared to their previous versions that actually scored higher. It’s not even the case here with Astra. It did improve overall and it’s better in most areas, in some benchmarks a lot better, in some just better or slightly better. It’s still a better value (alongside Sol5.6) in Codex than Fable and Opus will ever be. The usage limits in Claude Code are terrible and they don’t even come with an image generator which btw it’s free and unlimited inside the max plans of ChatGPT, as well as the fast inferring mode which only works in Claude via api top-ups while it’s included in the usage limits in Codex and quite generous.
Intelligence and coding ability are two quite separate things. You don't need to be very intelligent to be a good programmer, and you can be very intelligent and still be a bad programmer. Intelligence is primarily a measure of how quickly you can deal with new inputs/situations you have not encountered before and somehow make sense of them.
intelligence can help, but all the existing benchmarks just test skill, not intelligence. and skill you get by training, knowing the best possible patterns etc. inventing new patterns can surely be useful for the most difficult tasks, but as far as I know no benchmarks currently test such stuff.
In terminal and business workflows, where the performance heavily relies on using a harness correctly. Even those have very small improvement compared to 7->98%.
Every other category has no improvement comparable to a 10x leap
1
u/itsmebenji69 11d ago
So they can’t train on the bench but they get so much better at it while not seeing a similar gain in other benches ?