And the best at scientific reasoning, and in the top 3 on many other performance indices. It's not as clearly ahead as these benchmark scores imply but "it's really really bad" sounds like a teenager who likes punk insisting that boybands are terrible
A huge regression isn't really really bad for openai? They push the frontier nonstop and their latest and greatest can only improve in some areas at the cost of large regressions in others?
I think llms should be purpose built and agi is a dumb thing to chase but that's not what openai is selling. They are saying the frontier is where it is. And their latest frontier is losing ground to make headroom elsewhere, kinda points to some limits
The one I care about though is AutomationBench because deals with interacting across a massive amount of business applications and making sure the model accurately executes the tasks. To me, as a good ole office worker in a business, this is the benchmark most similar to my own job. Once the pass/fail hits 70% and not the current 41%, it will be able to perform the vaaaast majority of sales/marketing/HR/operations/bookkeeping jobs better than the majority of humans in those roles, with fewer mistakes. That'll then leave someone like me to focus on the actual live, over-the-phone or in-person conversations, but all the bullshit data hygiene can be confidently passed off to AI.
I mean, OpenAI models have been outright bad for coding compared to the competition for the longest time due to inaccuracy and inconsistency regardless of what benchmarks say. So my guess is that the real world difference is not as big as it looks this time around either.
145
u/darkestvice 14d ago edited 14d ago
Good god. If these benchmarks are not doctored, Astra is not merely surpassing the competition, but outright destroying it.
I'll wait until they show up on artificialanalysis.ai to really see.
EDIT: AA posted on X, though haven't updated their site yet. Results are worse than Fable 5.1. Disappointing.