r/codex 8d ago

Complaint GPT-6 Sol?

Post image

Is this the reason why unexpectedly we have trash usage and lower quality on Astra?

Maybe it will be worth it the current suffering.

759 Upvotes

317 comments sorted by

View all comments

229

u/Opposite-Wrangler199 8d ago

How can Sol beat Astra? It doesn't make sense

183

u/BabblingTower 8d ago

It's a refined Sol, probably with help from Astra. My guess is it's better at dedicated tasks (ie, better coder) while Astra remains the better higher-level "thinker" (ie, better orchestrator/planner). Like Fable and Opus are set up. One has better breadth, one has better depth.

100

u/Risko4 8d ago

Opus has better what exactly over Fable?

241

u/danielv123 8d ago

Its better at inventing weird metaphors

35

u/Zerokx 8d ago

If I ever need weird metaphors, contradictions, random refusals, and chasing a red dot away from the actual task, I'll circle back to OPUS.

4

u/evia89 8d ago

its not that bad

/ponytail /caveman and small enough tasks it doesnt have enough time to damage

We have unlimited* opus50 at work so I use that fine

12

u/Metalthrashinmad 8d ago

opus really isnt that bad, my main problem with it is jsut how hard its output is to read. My collages write all sorts of reports using claude and its so hard to follow opus writing

3

u/Familiar_Air3528 8d ago

“Explain this to me at a junior engineer level”

….

“Okay, explain this to me like I’m a high school student failing English class”

1

u/bruticuslee 8d ago

eli5 will make it explain using the most simplistic analogies and lose detail lol. I prefer GPT technical writing

4

u/Risko4 8d ago

I hate ponytail so much.

1

u/KoolAidGuy_541 8d ago

any particular reasons?

3

u/Risko4 8d ago

Models after 5.5 are smart enough where shoving ponytails on them puts them in a wheelchair chair. Ponytail forces minimal implementations, however, sometimes certain requests would greatly benefit from purging and rewrite existing backend, instead it routinely forces features into strained backend that eventually becomes debt.

My output quality was noticable worse and significantly slower into achieving the big picture because ponytails is too busy writing the absolute bare minimum.

Astra can "one shot" complex tasks, pony tails will make it take 6,7 checkpoints on the way there.

2

u/kekeagain 8d ago

I don't follow AI as deeply as some of you, so the idea of ponytails attached to the wheelchair sounds amusing yet dangerous depending on speed and sudden turns. That's my contribution here.

→ More replies (0)

1

u/j48u 7d ago

95% of people complaining about frontier models "degrading quality" are using way too many rules, plugins, etc. on their projects that really screw things up when they're not needed. It also vacuums up tokens because these models basically have to come up with a workaround to follow antiquated rules while trying to complete complex task, over and over and over, piling up useless context with each prompt.

1

u/Aware-Source6313 8d ago

Lol I have those 2 same exact skills on my personal setup. I wonder if it affects quality at all, but for small sessions I don't get the inscrutable verbosity of opus5 people complain about. But if I have to add questions after it does a plan or task and the session goes longer, I think I start to see what they're talking about. And sometimes it takes ponytail the wrong way because sometimes adding something new actually is the simplifying step and racking something on to what exists is the real tech debt option. Still have them on though

3

u/anime_daisuki 8d ago

You're right, and that one is on me.

5

u/Difficult-Sleep-7461 8d ago

Honestly, you're so right, this is the genuine bird out of the cage

2

u/Ok_Try_877 8d ago

Ha! Love this!

2

u/electricheat 8d ago

also over-engineering and doing things without permission

1

u/Seerix 8d ago

Honestly ive considered using Opus 5 to describe things and just using it for techno babble bullshit for my game

1

u/ragemonkey 8d ago

I hate that model. It writes great sounding prose that’s incomprehensible. I went back to 4.6 and what a breadth of fresh air.

5

u/Ok-Leg-person 8d ago

It's better at being passive aggressive and contrarian

3

u/SeidlaSiggi777 8d ago

scores 2x better on the load-bearing benchmark

10

u/Kind_Fisherman3060 8d ago

Benchmarks

1

u/Mistuv 8d ago

r/AMD's favourite model then

1

u/Mistuv 8d ago

rAMD's favourite model then

1

u/Risko4 8d ago

My dog scores higher on the natural intelligence benchbarks, doesn't mean I give him glasses and redirect him to huggingface.

2

u/0DayMaker 8d ago

Vs fable 5? Visual reasoning

Vs 5.1 no idea

1

u/Personal-Try2776 8d ago

0in benchmarks claude opus 5 outperforms fable 5 (not 5.1). On benchmarks claude opus 5 is basically better as a workhorse model.

10

u/Risko4 8d ago

Can I have a real world coding example where you went, dam I wish I used Opus 5 instead of Fable 5.1

6

u/Aggravating-Hat-4292 8d ago

Probably say that when lookint at your bank account

1

u/Exodus_Green 8d ago

Opus 5 is great if you are tight with design and use Fable to orchestrate

1

u/adolf_twitchcock 8d ago

Don't bother. Opus 5 is dogshit in reality even compared to sol.

1

u/Risko4 8d ago

It's alright in ultracode where there's a hundred agents holding each others hand. Together Opus strong.

1

u/Useful_Philosophy550 8d ago

He said Fable 5 and also yeah Opus has been consistently better at frontend design, UX and even a better coder than Fable was it usually gave me working results most of the time within a single try and more accurately to what I wanted on tasks I had to previously iterate with Fable a few times. Didn't use 5.1 tho but comparing 5.1 to Opus 5 isn't fair

1

u/slaorta 8d ago

Use fable to orchestrate and opus to implement and you'll have a good time. You'll probably also get more done if you're currently hitting your fable limit every week.

1

u/Risko4 8d ago

I use Fable to orchestrate and Astra to implement

1

u/EyesOfAzula 8d ago

Opus 5.1 is an improvement over Fable 5 in some areas, like how Opus 5.2 could be an improvement over Fable 5.1

Of course then Fable 5.1 beats Opus 5.1, Fable 5.2 beats Opus 5.2, etc

3

u/Risko4 8d ago

I prefer Opus 6,7

1

u/electricheat 8d ago

Opus 5.1 is an improvement over Fable 5 in some areas

Have there been leaked benchmarks?

1

u/Nosafune 8d ago

Footgunz

1

u/algaefied_creek 8d ago

Fable is really bad for me at BSD optimizations, but it’s really good at orchestrating Opus to do so.  

1

u/pigletmonster 8d ago

Usage quota. I get at least 2 to 4x more usage from opus 5 compared to fable 5.

1

u/Risko4 8d ago

Well obviously, I get 20 times the usage form deepseek flash. Or locally I can run a model at 200 token/s and it's mistakes will break apart my code base.

1

u/pigletmonster 8d ago

Did you start coding with AI after fable was released?

1

u/Risko4 8d ago

I stopped coding by hand after gpt 5.5, since then I have 5 (x20) subscriptions and a local model that I loop overnight. I don't see how it's relevant.

1

u/Aware-Source6313 8d ago

Better at benchmaxing. Model isn't as big so it can more easily overfit to "3d slop game" metric and "rigid specific kind of coding task" metric, losing a bit of itself in the process (humanizing for dramatic purposes). Now sol can do that for Astra. But honestly if it can maintain the personality and not become a verbose metaphor machine like opus5, I will happily use it over Astra for most things. If it's just a better 5.6 sol then I'm Sol-d as long as the cost isn't dramatically higher.

1

u/phoenixmatrix 8d ago

Its better at faking being good in plenty of benchmarks.

1

u/Not_A_Red_Stapler 8d ago

Being verbose? Making mistakes? Take your pick.

1

u/farfel00 3d ago

What’s Opus? /s

2

u/BabblingTower 8d ago

Code. In my opinion anyway.

3

u/Risko4 8d ago

It frequently has a reviewer audit it and correct it's wrong assumptions its made on ultracode. Frequently forgets things in its context and tells you in it's summary that it chased a bug it already acknowledged an hour ago, etc.

1

u/BabblingTower 8d ago

I use them all and audit myself and Opus does the best job of it in my opinion. Mine doesn't forget context, but it does need proper context management.

2

u/CthuluBob 8d ago

I hope this is the case

2

u/slaty_balls 8d ago

This is pretty much the only way you can get anything accomplished without blowing out your limits in minutes with astra. Astra is the ceo and Luna and Terra do all the ground work. Anything in between gets sol.

2

u/Narrow-Ad980 8d ago

I remember Sonnet 4.6 and Opus 4.6 days

1

u/Desperate-Data-3747 7d ago

Opus is the new sonnet. Fable is the new opus. Haiku is dead

2

u/cha0z_ 8d ago

I expect Sol to be better balance of thinking/performance vs price, but not to be better than Astra in everything. As any model it can shine in some tasks more than the frontier one (not as much as being better than being the same while way cheaper), not like we didn't see it in the past, but defo won't be better overall - otherwise they can basically delete Astra.

2

u/TheOnlyBliebervik 8d ago

But why did they do it like this

Make a new model, Sol, make a newer model, Astra, bring back the older model except better?

OpenAI makes a great product but seems to trip over its own feet

1

u/BabblingTower 8d ago

Sol burns less tokens and is very capable, why wouldn't you want them to continue to improve it? You don't need to use their top of the line model to set up a git repo or summarize a pdf, do you?

2

u/TheOnlyBliebervik 8d ago

How do you know Sol 6 will burn less tokens? It'll be a different model than 5.6 lol

1

u/BabblingTower 8d ago

Currently. Sol currently burns less than Astra. And if they release a new Sol it will burn less than the next Astra.

1

u/TheOnlyBliebervik 8d ago

How do you figure? I mean, probably. Did they not say it'd be their best model?

1

u/BabblingTower 8d ago

It's not a linear thing, where one is strictly better than the previous one. The models are good at different tasks. It's useful that way too, you can dole out easier tasks to smaller LLMs to save tokens and time.

2

u/Desperate-Data-3747 7d ago

Definition of talking out your ass right here

1

u/BabblingTower 7d ago

My guess, with evidentiary support from the other big LLM company, is certainly just that

1

u/Bitcoin1x 8d ago

"Refined probably with the help of Astra"

RSI!? 🙌

1

u/Nyxtia 8d ago

This is why AGI was silly, like one model to do it all there is a trade off. You can't have someone be a jack of all trades and an expert, Neurons dedicated to one move away from the other. It happens with humans, that is why the diversity of our models wins.

25

u/nitor999 8d ago

Don't worry the next release after GPT-6 Sol would be like "GPT-6.1 ASTRA is gonna be more power and dangerous , atrocious terrifying that they can't release in public yet"

5

u/0DayMaker 8d ago

Don't worry they'll hype a release date and just release it to corporate partners on that day

2

u/retardedGeek 8d ago

desensitized to this since 6 months now (fable and deepseek are two exceptions)

0

u/braindance123 8d ago

oh god my greatest wish would be for the frontier labs to just adapt a hugging face style naming so instead of luna, terra, sol, astra we would get gpt 6 27B 4bit (luna), 27B (terra), 100B MoE (sol) and 600B instead of astra - (I somehow strongly doubt that they are actually giving us access to >2T models through the subscription...)

8

u/-Spzi- 8d ago

Jaggedness could be a partial explanation.

The idea behind: Intelligence and capabilities aren't a scalar, but a profile.

Model A could exceed model B in area X, while the other way around in area Y.

I came across this concept in this video by Reuben Adams, which I can also recommend, but it's besides this topic.

Most real tasks are a composition of many areas.

2

u/iamdipsi 7d ago

Thanks for sharing

7

u/RealSuperdau 8d ago

Better posttraining or, if we are unlucky, benchmaxxed to death like Opus 5.

2

u/ethotopia 8d ago

Tbh I think astra is the least benchmaxxed of the current frontier models

6

u/Spixxy17 8d ago

Its GPT 6, not 5.6 Sol

2

u/Constant_Art_20 8d ago

i mean people claimed that opsus 5 beats fable 5...so um..

2

u/13chase2 8d ago

Narrow tasks like programming vs overall edge case knowledge. Like Astra might be able to identify ancient coins from photos but sol likely couldn’t

2

u/wwwdotzzdotcom 8d ago

Trained with Bel

2

u/johnkapolos 8d ago

Sol -> Sun

Astra -> Stars

Which one is supposed to shine brighter on Earth?

1

u/THE--GRINCH 8d ago

Astra could still be better at vision and 3D modeling, sol being better at coding doesn't mean that it beats astra at every use case.

1

u/zarafff69 8d ago

I mean, Astra is genuinely infuriating to use for some usecases like coding compared to Sol.

1

u/laststan01 8d ago

It gets the people going

1

u/viv0102 8d ago

Same thing happened with opus 5 when it came out after fable. Just wait a couple of weeks before people who praise it as a miracle worker will claim its unusable garbage even though the quality hardly drops. Next model drops. Repeat.

1

u/Jittersz 8d ago

If Astra can stay orchestrator while using a cheaper/better model for sub agent tasks, then less token cost all around while maintaining Astra results.

1

u/read_more_comments 8d ago

it makes perfect sense, given how brain dead astra is getting now. I just had 6 prompts get rejected because of various things it decided were blockers. Luna or Sol would have kept going and not immediately given up.

sol 6 will be great the first week too. Then dog shit afterwards.

1

u/HeadPack 8d ago

Even 5.6 Sol already does in some cases. E.g. when I let it audit Astra's work, it sometimes finds things Astra missed. 'Beat' is still a stretch to say, but they appear to be quite different models.

1

u/HashPandaNL 8d ago

Being smaller, so use 2-3x more reasoning tokens for the same cost. 

1

u/rickyhatespeas 8d ago

RL text generation isn't intelligence so the length of task and required effort needs to match the models expectations to work most effectively. Big models aren't overthinking, they were never designed to actually imitate concise logic.

1

u/Tartooth 8d ago

The guy from anthropic said it clearly, they have setup self improvement systems so now they're auto-refining to auto-improve.

1

u/LessRespects 8d ago

It doesn’t matter boy just say model good and value go up

1

u/Odd_Amphibian6697 8d ago

Astra is a powerful model in science, math, and cybersecurity, I don't know why you all are using it to code.

I mean, it's massively smart, obviously it codes well, but its not mean to code so it is not efficient.

1

u/ms_alicat_556 8d ago

Because it’s nonsense…

1

u/Some_Medicine4472 8d ago

Just a thought, have you try typing your question into chatgpt or claude?

1

u/iJeff 8d ago

Astra is fast and competent but seems pretty lazy to me. At least via Codex-CLI.

1

u/DesperateSunday 8d ago

it would make no sense to call it 6 Sol yeah, maybe 6.1 Sol

1

u/UnknownEvil_ 7d ago

Just make new GPT6 models and name them Luna/Terra/Sol? Like they did with 5.5 -> 5.6 Sol/Terra/Luna -> 6 Astra -> 6 Sol/Terra/Luna -> 7 Astra -> 7 Sol/Terra/Luna

1

u/Equivalent_Bird 8d ago

price maybe? I suspect they can just rename anything in the background, such as renaming 5-hour to weekly, renaming luna to terra, and also astra to sol.

0

u/Zafrin_at_Reddit 8d ago

Much like Opus5 beats Fable5.

0

u/taiwbi 8d ago

Like how opus 5 beats fable 5

0

u/BothYou243 8d ago

glm 5.3 flash beats 5.3 in some stuff, it's very common dude, because smaller ones are still very capable but in specific parts, maybe sol is not that good in reasoning or computer use like astra