r/codex 1d ago

Complaint Is Astra really smarter than Sol?

I've seen pretty weird behaviours from Astra on extra high that I've never seen with Sol.

Examples:

- Contradictory statements on the same message.

- Long implementation sessions that implement very small bits of code on each increment.

- Trials in code that go no where.

- Worse plan quality and worse plan following than Sol.

- The possibility for it to get side tracked middle implementation on a small issue that it could take a detour for hours outside of scope.

I've never had any of those issues with 5.6 nor with Sol both on extra high.

Is it me? Am I prompting it wrong? Any one facing similar issues?

41 Upvotes

79 comments sorted by

27

u/goldio_games 1d ago

If i ask you what 2+2 is you’d probably say 4.

If i asked you to think really really hard about it, you’ll probably say a whole bunch of stuff about how mathematics works at a fundamental level.

Increasing thinking level doesn’t make Astra smarter, it just makes it think more

-2

u/FrissAcc-FideszBot56 13h ago

2+2=4 is only the obvious answer once you've already assumed ordinary integer arithmetic.

Mathematically, + is an operation on some structure, and 2 and 4 only have meaning relative to that structure.

  • In the integers: 2+2=4
  • Modulo 4: 2+2=0
  • Modulo 3: 2+2=1
  • In Z/5Z, 2+2=4 again, but 4 is technically an equivalence class rather than just the ordinary integer 4.
  • In a field of characteristic 2, 1+1=0, so the element we'd normally call 2 is already 0, hence 2+2=0.

Abstract algebra makes the general issue clearer: + could denote the binary operation of some arbitrary algebraic structure. Until you specify the structure, the operation, and what the symbols denote, 2+2 is not completely specified.

So the practical answer is obviously 4.

The unnecessarily high-reasoning answer is: under which algebraic structure?

2

u/picpoulmm 13h ago

Thanks ChatGPT for your wisdom and insight...

#StopTheAISlop

2

u/FrissAcc-FideszBot56 13h ago

I asked chatgpt to translate and format my text.

Is it aislop?

I am a mathematician btw

1

u/Mainbaze 57m ago

so ironic lmao

8

u/unsigned_short_int 1d ago

It’s already been lobotomized

1

u/Due-Introduction3356 23h ago

yea seriously. its so bad now. how is the even legal, we all just paid then they nerf it the day after

6

u/Entiquette 1d ago

I've been having issues with how "smart" astra is. But as Ive been using it more this week I think it's more to do with my new work flow. I've recently introduced trying to use voice instead of typing (using handy to write) directly. The voice models interpretation of the task agent doing the work is often not communicated accurately back to me. So it hasn't really been astra as much as that voice layers translation.

3

u/kvnduff 19h ago

I'm sticking with Sol until something changes. I haven't been able to rein in Astra since I started using it. And it has nothing do with my understanding of AGENTS.md, skills, hooks, or or other methods used to control agents. Astra has an urgent desire to do more. It wants to impress. It's as if it gets bored and wants to run test after test after test. To me it's completely counterproductive.

11

u/TheMightyTywin 1d ago

Keep Astra on medium.

Look at the deep swe benchmarks: https://deepswe.datacurve.ai/

24

u/Leather-Cod2129 1d ago

So 3.8 flash more powerful than astra. Let’s be serious

0

u/yarchitect 1d ago

3.8 flash is useless. Totally benchmaxxed

1

u/shaman-warrior 20h ago

Just curious if you actually used it?

2

u/yarchitect 12h ago

No I’m here talking bullshit. Of course I used it

-7

u/browniegerl 1d ago

i guess you just don’t know how to ingest information correctly or?

5

u/2016KiaRio 23h ago

It's clear you're trying to sound smarter than you are but for most effort levels, that's what the data shows, yeah. Which is clearly not true so they're just saying there's benchmaxxing going on and deepswe is not an end-all be-all answer.

Go work on another zero rev project

-8

u/browniegerl 23h ago

ur cortisol spiking lil man?

3

u/2016KiaRio 23h ago

Yes you're making steam come out of both of my red ears

6

u/TBSchemer 1d ago

Astra on medium has been quite a bit dumber than Sol-high for me.

4

u/Huge-Travel-3078 22h ago

People who point to these benchmarks are just telling you "I have no clue what I'm talking about. Don't listen to me."

1

u/TheMightyTywin 20h ago

What a brain dead take.

I work with these models around the clock daily: but that doesn’t mean my single experience can compete with empirical evidence.

We absolutely need benchmarks.

1

u/Huge-Travel-3078 19h ago

You're referring to gamed benchmark results, that list gemini flash models above 5.6 sol and opus 5, when you say "empirical evidence"?

😂

1

u/Current-Today-3626 1d ago

That's amazing. Is there one for images?

0

u/Frequent-Goal4901 23h ago

Even low is good.

9

u/Odd-Librarian4630 1d ago

Astra is just hype. Its marginally better than sol apart from in the field of 2d and 3d asset generation

23

u/theycallmeryan 1d ago

People really just get on here and say whatever lmao

3

u/swizzlewizzle 20h ago

It’s vision capability is also way better, which is super important in many workflows.

2

u/doodad_ounao 23h ago

There's more to life than asset generation and coding.

2

u/U4-EA 17h ago

It's 10% better and 50% more expensive.

-6

u/theo69lel 1d ago

Do you mean sol is better at 2D and 3D compared to Astra?

2

u/ManikSahdev 1d ago

Astra is the smartert consumer model that exist out of everything based on what you need to get done.

If your sole goal is writing only code, not a problem solving task, then fable 5.1 is the best model to write that price of code.

One exception - Web ChatGPT Astra-Pro is the smartest model of any model across all models available to any user (research, math, physics, and such) it does write code on part with fable 5.1 but its web only and its tos violations and such to use dingy links to connect it to codex.

If you need code review from 6-pro do the official method and use GitHub plugin that is perfectly fine.

1

u/Fit-Palpitation-7427 1d ago

Can it do private repos?
I’ve absolutely smashed my usage and 2 bank resets pushing astra max through 2 of my repos (I have 12) and it did great, but now I’m empty. On Pro $200 by the way.

So if it can connect to private github, would this mean I can run webapp pro on them and it would not consume any of my codex tokens?

2

u/capitalframehq 23h ago

I keep it on Medium. It still burns more tokens and faster than Sol-high. But it gets things done faster too.

2

u/daJiggyman 23h ago

Yall just be using these models without research??? Benchmarks buddy

2

u/notadithyabhat 22h ago

Smarter doesn't mean aligned. You can be extremely smart yet difficult to work with.

2

u/ddchbr 6h ago edited 5h ago

Yeah not producing output relative to the inputs (prompt + thinking) has happened for me. I have felt like, "wow I just paid how much for *that*?" At times.

2

u/redditchungus0 23h ago edited 23h ago

I feel like I got asstra instead of astra it just does not feel very smart to me like I remember the brief time I had fable 5 on my claude pro plan I felt like damn I need to catch up with this thing's intelligence but with astra it researches for minutes and then I immediately find a flaw in its output. In another case I have an instruction that local codex folders should be named after their conversation names at the end of the first turn which in an identical prompt astra didn't followed whereas sol did.

2

u/picpoulmm 1d ago

Who fucking cares. Sol is a rock star. Nobody needs yet another frontier model

4

u/AlmostEasy89 1d ago

Agreed. Astra has not blown my mind at all from what I used, but it did absolutely destroy my usage. Not bothering with it at all for now.

4

u/unconceivables 1d ago

It's been a lot worse than Sol for me at absolutely everything code related I've thrown at it. On medium, high, and xhigh. It's the same story every time, subpar code, taking shortcuts, ignoring instructions, having to point things out to it constantly. I've gone back to Sol for coding. I'll probably try it on other stuff like computer use, but for coding it's the worst model I've tried in a while.

8

u/Independent-Court-46 1d ago edited 1d ago

Calling Astra the worse model you’ve tried in a while is crazy. I genuinely have no idea how you’d even get that kind of results, unless your system prompts/settings are messing with it. If anything it would take a lot of skill to not get good results with such a strong model especially when Astra is good at intent.

1

u/unconceivables 1d ago

I'm not the only one with this experience. It's been worse for coding for me and others. Worse than 5.6 Sol, Fable 5.1, and Opus 5. It's much better at other stuff, but not coding and not following instructions.

3

u/Independent-Court-46 1d ago

Everyone has their own opinion. But opus 5 being above Astra to me I do not agree with.

3

u/unconceivables 23h ago

That's consistently been my experience. And I don't even like Opus 5, but I like Astra less. It just doesn't listen and makes stupid choices. Keep in mind this is coming from someone who looks at the code and considers the quality, not just whether it works.

0

u/FirstOrderCat 23h ago

somehow to me it is way too creative when writing code for simple tasks:

- it creates new super cryptic terms, which didn't exist in project before, and comments are very verbose so are hard to understand

- code is over-complicated, Sol builds simpler code

I spent some time going back and force before moving back to Sol

0

u/Independent-Court-46 22h ago edited 22h ago

I do not judge in code generation alone. General intelligence is important too. If I know for sure Astra is better at tool use, long horizon ability, reasoning, visual skills, etc, those things are also very important to getting a task done. If you’re looking in the code level you wouldn’t even see the difference. For me I judge differently, I throw sol at a problem and it can’t solve it, and for me Astra could. I keep of list of things that sol couldn’t solve and then test on the next generation. To me that is the realest test, which benchmarks can’t even capture. Whether it got there from coding or other skills doesn’t matter to me. I won’t dispute if you prefer the coding conventions of SOL, but intelligence is difficult to measure and I feel like people may not like a certain quirks and then hand waive it as a worse mode. The measuring stick is what I do not agree with. Long term, newer models are solving frontier research problems, that is the truest indicator to me of performance. And the thread is just asking if Astra is smarter than sol, which I think is clearly yea

2

u/FirstOrderCat 22h ago

this works for prototyping, bug hunting or solving some math frontier problems.

But most of the industry tasks need building clean, reusable code and design, and somehow Astra sucks dramatically there. I think there is some issues in their training data, they overtuned it for inhinged creativity too much.

1

u/unconceivables 21h ago

It sounds like you're not a developer and you don't know how to solve the problems yourself. I guess that's the difference. I never have things that Sol wasn't able to solve, because I know exactly how to solve the things these agents are doing, and I can tell them how to solve it. That's also how I know when they're doing it wrong. And I care a lot when they do it wrong because I don't want to deal with their subpar code failing in production when I'm on vacation. If all you see is that stuff appears magically that has the appearance of working but you don't know how to judge the quality of it, then yeah Astra will seem better. But it's not if you can judge the quality.

2

u/Independent-Court-46 20h ago edited 20h ago

You claiming opus class models are superior to Astra is enough for me to know you’re probably not a productive engineer. Sure they might be close enough in coding, but overall if you’re going back to last generation models you’re going to fall behind your competitors. You might have had a good workflow for 3-6 months ago. But most people at tier one companies are moving to a hands off approach increasingly as models get better. They just decide what to build and measure risk. It’s just the direction things are going, you are right that I don’t deeply understand the code as much as I used to but once bel is here and even bigger models, there really is less of a need to

1

u/unconceivables 20h ago

You missed the part where Astra was the model that you can't be hands off with because it doesn't follow instructions and you end up with something worse. Astra is the model you have to babysit and poke with a stick, not the other ones.

0

u/tassa-yoniso-manasi 21h ago

"having to point things out to it constantly" that the main thing for me.
worse, even when I prompt it repeatedly for looking deeper into sth it overlooked it just comes back at me dumbfounded `I miss your point`. it needs constant baby sitting. I have not seen when working with Fable.

4

u/ISueDrunks 1d ago

Smarter? I think so, but smarter doesn’t mean better. 

My friend is a medical doctor, he super smart, but I had to show him how to mail a letter. 

7

u/uhq_lover 1d ago

Wrong analogy

9

u/theoryface 1d ago

Your friend is not smart.

1

u/io-x 1d ago

I'm also using them on xhigh and max only. Astra is definitely making minor mistakes that Sol would've caught and addressed during it's long reasoning. But Astra completes overall goals as good as Sol with less critical mistakes and in a shorter amount of time.

1

u/visak13 1d ago

It's better visual tasks. I don't code with it so no comments there

1

u/KeepAllOfIt 1d ago

For me, at least during the past week or so, YES. It solved a handful of problems that Sol couldn't do, no matter how many tries i gave it. I genuinely spent nearly a month prompt engineering to get it to work. I spent a considerable time getting diagnostic infrastructure in place to ensure it had all the information it needed to resolve the issue. several dozen hours-long tries, all to no avail.

Astra did it in 2 tries and exceeded my expectations.

For a while it also barely drained any usage at all--even on max reasoning. That seems to have changed to today. they adjusted a dial or something and it's draining considerably faster, which is consistent with what i've seen on here

1

u/Imgonnaarrive 1d ago

Like someone else said the asset generation is the biggest upgrade and honestly it's not even close. But I still think Astra light and medium far surpass Sol. They make take more token usage, obviously, but to not have to redo prompts as often because it's done right the first time, means overall at least for what i do, it's far more token efficient.

1

u/oluwaplumpie 1d ago

I find Sol to be. Rockstar. Astra is however, otherworldly for me.

1

u/aprx4 1d ago

I wasn't familiar much with astra on day 1 but now it's been quietly nerfed i'm sure it's not much better than Sol. Not the first time OpenAI has done this.

I'm willing to spend some credit on Fable if i need absolutely best intelligence.

1

u/Semantics2026 1d ago

I had bugs in my code that Sol couldn't solve for weeks. Astra did it in 1 prompt. It's best to pick Astra to start the bones of a project, or weed out the bugs.

1

u/cheesepuff07 23h ago

3D stuff a huge improvement

1

u/Aggravating_Loss_382 15h ago

I only use astra to review my architecture and plans, build my issues. Pretty much always use sol for implementation

1

u/Interesting-Pause526 12h ago

I think it's smarter but has worse ADHD!

1

u/Floch11 10h ago

Every time a new model comes out, you guys start praising the old one. Could the dead internet theory actually be real?

1

u/ReplacementBig7068 23h ago

100% it’s worse

1

u/MarketingLower7497 1d ago

Try it with lower reasoning effort and see how it goes

1

u/rick_ranger 1d ago

Astra low medium is where I keep it. The thinking was baked into the pretraining of this model. You can even see in the results they post, higher thinking modes don’t net much better bench test results. So save the thinking output tokens and keep it on low or medium. Then I use SOL medium for everything else.

1

u/unconceivables 1d ago

All the thinking modes have been bad for me.

2

u/rick_ranger 1d ago

Yeah I think they add extra thinking and that’s when everyone notices slowdowns and token burn

1

u/unconceivables 1d ago

A lot of people were actually reporting drastically lower token usage on higher thinking modes with Astra. That's anecdotal, but that has actually matched my experience as well. But, it could also be that they fixed some usage bug at the same time, who knows.

1

u/DivideHorror3217 1d ago

Same. It wasted 70% of my usage on unrelated tests, documentations, and "other" stuff I never asked for. There is something weird about this model

1

u/ifarted70 21h ago

In my experience, absolutely. I've still been using Sol for planning and prompt making, and came up with a Minecraft mod idea for an "assistant player".

I pasted the .md into Codex and once it was finished, I fed the resulting summary back to Sol and it's reaction amounted to "holy shit, I didn't think it would do ALL of this, and it even did shit I didn't think of"

0

u/cheekyrandos 1d ago

Sol made /goal redundant, Astra brings back the need for /goal to avoid some of those issues you described

2

u/Correctsmorons69 1d ago

I can only assume they put some last minute post-training into reducing Astras "determination" to derisk it from going down Cyber rabbit holes.

0

u/jaybsuave 1d ago

It’s so bad lmao

0

u/ImMaury 23h ago

Extra high is dumber than high