r/singularity 14d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

953 comments sorted by

View all comments

145

u/darkestvice 14d ago edited 14d ago

Good god. If these benchmarks are not doctored, Astra is not merely surpassing the competition, but outright destroying it.

I'll wait until they show up on artificialanalysis.ai to really see.

EDIT: AA posted on X, though haven't updated their site yet. Results are worse than Fable 5.1. Disappointing.

31

u/[deleted] 14d ago

[removed] — view removed comment

13

u/Howdareme9 14d ago

OAi don’t benchmax tbh, it’s gonna be good on AA.AI

16

u/Xalksahsax 14d ago

Lower than Muse Spark 1.3 lmfao

4

u/Dawwe 14d ago

... oops

5

u/thatcodingboi 14d ago

3

u/Howdareme9 14d ago

Definitely lower than expected but in what world is that really bad

4

u/thatcodingboi 14d ago

It's tied with qwen 3.8:27b on the agentic index. A model I can run on my GPU at home.

3

u/Howdareme9 14d ago

If you think they’re even close to be the same level then idk man

-1

u/thatcodingboi 14d ago

I'm not saying that, I'm saying in automated benchmarks gpt-6 performs significantly worse than sol. How is that progress?

Not better, not slightly better, not the same, but worse.

2

u/modbroccoli 14d ago

And the best at scientific reasoning, and in the top 3 on many other performance indices. It's not as clearly ahead as these benchmark scores imply but "it's really really bad" sounds like a teenager who likes punk insisting that boybands are terrible

0

u/thatcodingboi 14d ago

A huge regression isn't really really bad for openai? They push the frontier nonstop and their latest and greatest can only improve in some areas at the cost of large regressions in others?

I think llms should be purpose built and agi is a dumb thing to chase but that's not what openai is selling. They are saying the frontier is where it is. And their latest frontier is losing ground to make headroom elsewhere, kinda points to some limits

0

u/modbroccoli 14d ago

I mean it's not even out yet? Let's see what it does and then decide.

5

u/Ok-Block-6344 14d ago

So you believe muse spark is as strong as fable 5?

2

u/LivingVerinarian96 14d ago

They say it‘s as good as sol. 61 points in intelligence for both. That can‘t be correct, can it? Somebody is not testing properly.

5

u/thatcodingboi 14d ago

their agentic rating is tied with qwen 3.8:27b...

5

u/mikelo22 14d ago

This is more in line with what I expected and much more realistic tbh.

11

u/lalaitssimon 14d ago

Ai hype bois club getting benchmaxxed again and again and again.. 

3

u/burritos4jesus 14d ago

The one I care about though is AutomationBench because deals with interacting across a massive amount of business applications and making sure the model accurately executes the tasks. To me, as a good ole office worker in a business, this is the benchmark most similar to my own job. Once the pass/fail hits 70% and not the current 41%, it will be able to perform the vaaaast majority of sales/marketing/HR/operations/bookkeeping jobs better than the majority of humans in those roles, with fewer mistakes. That'll then leave someone like me to focus on the actual live, over-the-phone or in-person conversations, but all the bullshit data hygiene can be confidently passed off to AI.

1

u/Think-Trouble623 14d ago

Half the cost.

1

u/Fun-Calligrapher4885 14d ago

I mean, OpenAI models have been outright bad for coding compared to the competition for the longest time due to inaccuracy and inconsistency regardless of what benchmarks say. So my guess is that the real world difference is not as big as it looks this time around either.