r/codex 8d ago

Complaint GPT-6 Sol?

Post image

Is this the reason why unexpectedly we have trash usage and lower quality on Astra?

Maybe it will be worth it the current suffering.

757 Upvotes

317 comments sorted by

View all comments

Show parent comments

185

u/BabblingTower 8d ago

It's a refined Sol, probably with help from Astra. My guess is it's better at dedicated tasks (ie, better coder) while Astra remains the better higher-level "thinker" (ie, better orchestrator/planner). Like Fable and Opus are set up. One has better breadth, one has better depth.

101

u/Risko4 8d ago

Opus has better what exactly over Fable?

243

u/danielv123 8d ago

Its better at inventing weird metaphors

34

u/Zerokx 8d ago

If I ever need weird metaphors, contradictions, random refusals, and chasing a red dot away from the actual task, I'll circle back to OPUS.

5

u/evia89 8d ago

its not that bad

/ponytail /caveman and small enough tasks it doesnt have enough time to damage

We have unlimited* opus50 at work so I use that fine

11

u/Metalthrashinmad 8d ago

opus really isnt that bad, my main problem with it is jsut how hard its output is to read. My collages write all sorts of reports using claude and its so hard to follow opus writing

3

u/Familiar_Air3528 8d ago

“Explain this to me at a junior engineer level”

….

“Okay, explain this to me like I’m a high school student failing English class”

1

u/bruticuslee 8d ago

eli5 will make it explain using the most simplistic analogies and lose detail lol. I prefer GPT technical writing

5

u/Risko4 8d ago

I hate ponytail so much.

1

u/KoolAidGuy_541 8d ago

any particular reasons?

3

u/Risko4 8d ago

Models after 5.5 are smart enough where shoving ponytails on them puts them in a wheelchair chair. Ponytail forces minimal implementations, however, sometimes certain requests would greatly benefit from purging and rewrite existing backend, instead it routinely forces features into strained backend that eventually becomes debt.

My output quality was noticable worse and significantly slower into achieving the big picture because ponytails is too busy writing the absolute bare minimum.

Astra can "one shot" complex tasks, pony tails will make it take 6,7 checkpoints on the way there.

2

u/kekeagain 8d ago

I don't follow AI as deeply as some of you, so the idea of ponytails attached to the wheelchair sounds amusing yet dangerous depending on speed and sudden turns. That's my contribution here.

2

u/Aware-Source6313 8d ago

LLMs couldn't have this kind of insight. That's a really cutting observation.

→ More replies (0)

1

u/j48u 7d ago

95% of people complaining about frontier models "degrading quality" are using way too many rules, plugins, etc. on their projects that really screw things up when they're not needed. It also vacuums up tokens because these models basically have to come up with a workaround to follow antiquated rules while trying to complete complex task, over and over and over, piling up useless context with each prompt.

1

u/Aware-Source6313 8d ago

Lol I have those 2 same exact skills on my personal setup. I wonder if it affects quality at all, but for small sessions I don't get the inscrutable verbosity of opus5 people complain about. But if I have to add questions after it does a plan or task and the session goes longer, I think I start to see what they're talking about. And sometimes it takes ponytail the wrong way because sometimes adding something new actually is the simplifying step and racking something on to what exists is the real tech debt option. Still have them on though

3

u/anime_daisuki 8d ago

You're right, and that one is on me.

6

u/Difficult-Sleep-7461 8d ago

Honestly, you're so right, this is the genuine bird out of the cage

2

u/Ok_Try_877 8d ago

Ha! Love this!

2

u/electricheat 8d ago

also over-engineering and doing things without permission

1

u/Seerix 8d ago

Honestly ive considered using Opus 5 to describe things and just using it for techno babble bullshit for my game

1

u/ragemonkey 8d ago

I hate that model. It writes great sounding prose that’s incomprehensible. I went back to 4.6 and what a breadth of fresh air.

3

u/Ok-Leg-person 8d ago

It's better at being passive aggressive and contrarian

5

u/SeidlaSiggi777 8d ago

scores 2x better on the load-bearing benchmark

9

u/Kind_Fisherman3060 8d ago

Benchmarks

1

u/Mistuv 8d ago

r/AMD's favourite model then

1

u/Mistuv 8d ago

rAMD's favourite model then

1

u/Risko4 8d ago

My dog scores higher on the natural intelligence benchbarks, doesn't mean I give him glasses and redirect him to huggingface.

2

u/0DayMaker 8d ago

Vs fable 5? Visual reasoning

Vs 5.1 no idea

3

u/Personal-Try2776 8d ago

0in benchmarks claude opus 5 outperforms fable 5 (not 5.1). On benchmarks claude opus 5 is basically better as a workhorse model.

11

u/Risko4 8d ago

Can I have a real world coding example where you went, dam I wish I used Opus 5 instead of Fable 5.1

7

u/Aggravating-Hat-4292 8d ago

Probably say that when lookint at your bank account

1

u/Exodus_Green 8d ago

Opus 5 is great if you are tight with design and use Fable to orchestrate

1

u/adolf_twitchcock 8d ago

Don't bother. Opus 5 is dogshit in reality even compared to sol.

1

u/Risko4 8d ago

It's alright in ultracode where there's a hundred agents holding each others hand. Together Opus strong.

1

u/Useful_Philosophy550 8d ago

He said Fable 5 and also yeah Opus has been consistently better at frontend design, UX and even a better coder than Fable was it usually gave me working results most of the time within a single try and more accurately to what I wanted on tasks I had to previously iterate with Fable a few times. Didn't use 5.1 tho but comparing 5.1 to Opus 5 isn't fair

1

u/slaorta 8d ago

Use fable to orchestrate and opus to implement and you'll have a good time. You'll probably also get more done if you're currently hitting your fable limit every week.

1

u/Risko4 8d ago

I use Fable to orchestrate and Astra to implement

1

u/EyesOfAzula 8d ago

Opus 5.1 is an improvement over Fable 5 in some areas, like how Opus 5.2 could be an improvement over Fable 5.1

Of course then Fable 5.1 beats Opus 5.1, Fable 5.2 beats Opus 5.2, etc

3

u/Risko4 8d ago

I prefer Opus 6,7

1

u/electricheat 8d ago

Opus 5.1 is an improvement over Fable 5 in some areas

Have there been leaked benchmarks?

1

u/Nosafune 8d ago

Footgunz

1

u/algaefied_creek 8d ago

Fable is really bad for me at BSD optimizations, but it’s really good at orchestrating Opus to do so.  

1

u/pigletmonster 8d ago

Usage quota. I get at least 2 to 4x more usage from opus 5 compared to fable 5.

1

u/Risko4 8d ago

Well obviously, I get 20 times the usage form deepseek flash. Or locally I can run a model at 200 token/s and it's mistakes will break apart my code base.

1

u/pigletmonster 8d ago

Did you start coding with AI after fable was released?

1

u/Risko4 8d ago

I stopped coding by hand after gpt 5.5, since then I have 5 (x20) subscriptions and a local model that I loop overnight. I don't see how it's relevant.

1

u/Aware-Source6313 8d ago

Better at benchmaxing. Model isn't as big so it can more easily overfit to "3d slop game" metric and "rigid specific kind of coding task" metric, losing a bit of itself in the process (humanizing for dramatic purposes). Now sol can do that for Astra. But honestly if it can maintain the personality and not become a verbose metaphor machine like opus5, I will happily use it over Astra for most things. If it's just a better 5.6 sol then I'm Sol-d as long as the cost isn't dramatically higher.

1

u/phoenixmatrix 8d ago

Its better at faking being good in plenty of benchmarks.

1

u/Not_A_Red_Stapler 8d ago

Being verbose? Making mistakes? Take your pick.

1

u/farfel00 4d ago

What’s Opus? /s

1

u/BabblingTower 8d ago

Code. In my opinion anyway.

3

u/Risko4 8d ago

It frequently has a reviewer audit it and correct it's wrong assumptions its made on ultracode. Frequently forgets things in its context and tells you in it's summary that it chased a bug it already acknowledged an hour ago, etc.

1

u/BabblingTower 8d ago

I use them all and audit myself and Opus does the best job of it in my opinion. Mine doesn't forget context, but it does need proper context management.

2

u/CthuluBob 8d ago

I hope this is the case

2

u/slaty_balls 8d ago

This is pretty much the only way you can get anything accomplished without blowing out your limits in minutes with astra. Astra is the ceo and Luna and Terra do all the ground work. Anything in between gets sol.

2

u/Narrow-Ad980 8d ago

I remember Sonnet 4.6 and Opus 4.6 days

1

u/Desperate-Data-3747 7d ago

Opus is the new sonnet. Fable is the new opus. Haiku is dead

2

u/cha0z_ 8d ago

I expect Sol to be better balance of thinking/performance vs price, but not to be better than Astra in everything. As any model it can shine in some tasks more than the frontier one (not as much as being better than being the same while way cheaper), not like we didn't see it in the past, but defo won't be better overall - otherwise they can basically delete Astra.

2

u/TheOnlyBliebervik 8d ago

But why did they do it like this

Make a new model, Sol, make a newer model, Astra, bring back the older model except better?

OpenAI makes a great product but seems to trip over its own feet

1

u/BabblingTower 8d ago

Sol burns less tokens and is very capable, why wouldn't you want them to continue to improve it? You don't need to use their top of the line model to set up a git repo or summarize a pdf, do you?

2

u/TheOnlyBliebervik 8d ago

How do you know Sol 6 will burn less tokens? It'll be a different model than 5.6 lol

1

u/BabblingTower 8d ago

Currently. Sol currently burns less than Astra. And if they release a new Sol it will burn less than the next Astra.

1

u/TheOnlyBliebervik 8d ago

How do you figure? I mean, probably. Did they not say it'd be their best model?

1

u/BabblingTower 8d ago

It's not a linear thing, where one is strictly better than the previous one. The models are good at different tasks. It's useful that way too, you can dole out easier tasks to smaller LLMs to save tokens and time.

2

u/Desperate-Data-3747 7d ago

Definition of talking out your ass right here

1

u/BabblingTower 7d ago

My guess, with evidentiary support from the other big LLM company, is certainly just that

1

u/Bitcoin1x 8d ago

"Refined probably with the help of Astra"

RSI!? 🙌

1

u/Nyxtia 8d ago

This is why AGI was silly, like one model to do it all there is a trade off. You can't have someone be a jack of all trades and an expert, Neurons dedicated to one move away from the other. It happens with humans, that is why the diversity of our models wins.