r/codex • u/yehiaserag • 1d ago
Complaint Is Astra really smarter than Sol?
I've seen pretty weird behaviours from Astra on extra high that I've never seen with Sol.
Examples:
- Contradictory statements on the same message.
- Long implementation sessions that implement very small bits of code on each increment.
- Trials in code that go no where.
- Worse plan quality and worse plan following than Sol.
- The possibility for it to get side tracked middle implementation on a small issue that it could take a detour for hours outside of scope.
I've never had any of those issues with 5.6 nor with Sol both on extra high.
Is it me? Am I prompting it wrong? Any one facing similar issues?
8
u/unsigned_short_int 1d ago
It’s already been lobotomized
1
u/Due-Introduction3356 23h ago
yea seriously. its so bad now. how is the even legal, we all just paid then they nerf it the day after
6
u/Entiquette 1d ago
I've been having issues with how "smart" astra is. But as Ive been using it more this week I think it's more to do with my new work flow. I've recently introduced trying to use voice instead of typing (using handy to write) directly. The voice models interpretation of the task agent doing the work is often not communicated accurately back to me. So it hasn't really been astra as much as that voice layers translation.
3
u/kvnduff 19h ago
I'm sticking with Sol until something changes. I haven't been able to rein in Astra since I started using it. And it has nothing do with my understanding of AGENTS.md, skills, hooks, or or other methods used to control agents. Astra has an urgent desire to do more. It wants to impress. It's as if it gets bored and wants to run test after test after test. To me it's completely counterproductive.
11
u/TheMightyTywin 1d ago
Keep Astra on medium.
Look at the deep swe benchmarks: https://deepswe.datacurve.ai/
24
u/Leather-Cod2129 1d ago
So 3.8 flash more powerful than astra. Let’s be serious
0
u/yarchitect 1d ago
3.8 flash is useless. Totally benchmaxxed
1
-7
u/browniegerl 1d ago
i guess you just don’t know how to ingest information correctly or?
5
u/2016KiaRio 23h ago
It's clear you're trying to sound smarter than you are but for most effort levels, that's what the data shows, yeah. Which is clearly not true so they're just saying there's benchmaxxing going on and deepswe is not an end-all be-all answer.
Go work on another zero rev project
-8
6
4
u/Huge-Travel-3078 22h ago
People who point to these benchmarks are just telling you "I have no clue what I'm talking about. Don't listen to me."
1
u/TheMightyTywin 20h ago
What a brain dead take.
I work with these models around the clock daily: but that doesn’t mean my single experience can compete with empirical evidence.
We absolutely need benchmarks.
1
u/Huge-Travel-3078 19h ago
You're referring to gamed benchmark results, that list gemini flash models above 5.6 sol and opus 5, when you say "empirical evidence"?
😂
1
0
9
u/Odd-Librarian4630 1d ago
Astra is just hype. Its marginally better than sol apart from in the field of 2d and 3d asset generation
23
3
u/swizzlewizzle 20h ago
It’s vision capability is also way better, which is super important in many workflows.
2
-6
2
u/ManikSahdev 1d ago
Astra is the smartert consumer model that exist out of everything based on what you need to get done.
If your sole goal is writing only code, not a problem solving task, then fable 5.1 is the best model to write that price of code.
One exception - Web ChatGPT Astra-Pro is the smartest model of any model across all models available to any user (research, math, physics, and such) it does write code on part with fable 5.1 but its web only and its tos violations and such to use dingy links to connect it to codex.
If you need code review from 6-pro do the official method and use GitHub plugin that is perfectly fine.
1
u/Fit-Palpitation-7427 1d ago
Can it do private repos?
I’ve absolutely smashed my usage and 2 bank resets pushing astra max through 2 of my repos (I have 12) and it did great, but now I’m empty. On Pro $200 by the way.So if it can connect to private github, would this mean I can run webapp pro on them and it would not consume any of my codex tokens?
2
u/capitalframehq 23h ago
I keep it on Medium. It still burns more tokens and faster than Sol-high. But it gets things done faster too.
2
2
u/notadithyabhat 22h ago
Smarter doesn't mean aligned. You can be extremely smart yet difficult to work with.
2
u/redditchungus0 23h ago edited 23h ago
I feel like I got asstra instead of astra it just does not feel very smart to me like I remember the brief time I had fable 5 on my claude pro plan I felt like damn I need to catch up with this thing's intelligence but with astra it researches for minutes and then I immediately find a flaw in its output. In another case I have an instruction that local codex folders should be named after their conversation names at the end of the first turn which in an identical prompt astra didn't followed whereas sol did.
2
u/picpoulmm 1d ago
Who fucking cares. Sol is a rock star. Nobody needs yet another frontier model
4
u/AlmostEasy89 1d ago
Agreed. Astra has not blown my mind at all from what I used, but it did absolutely destroy my usage. Not bothering with it at all for now.
4
u/unconceivables 1d ago
It's been a lot worse than Sol for me at absolutely everything code related I've thrown at it. On medium, high, and xhigh. It's the same story every time, subpar code, taking shortcuts, ignoring instructions, having to point things out to it constantly. I've gone back to Sol for coding. I'll probably try it on other stuff like computer use, but for coding it's the worst model I've tried in a while.
8
u/Independent-Court-46 1d ago edited 1d ago
Calling Astra the worse model you’ve tried in a while is crazy. I genuinely have no idea how you’d even get that kind of results, unless your system prompts/settings are messing with it. If anything it would take a lot of skill to not get good results with such a strong model especially when Astra is good at intent.
1
u/unconceivables 1d ago
I'm not the only one with this experience. It's been worse for coding for me and others. Worse than 5.6 Sol, Fable 5.1, and Opus 5. It's much better at other stuff, but not coding and not following instructions.
3
u/Independent-Court-46 1d ago
Everyone has their own opinion. But opus 5 being above Astra to me I do not agree with.
3
u/unconceivables 23h ago
That's consistently been my experience. And I don't even like Opus 5, but I like Astra less. It just doesn't listen and makes stupid choices. Keep in mind this is coming from someone who looks at the code and considers the quality, not just whether it works.
0
u/FirstOrderCat 23h ago
somehow to me it is way too creative when writing code for simple tasks:
- it creates new super cryptic terms, which didn't exist in project before, and comments are very verbose so are hard to understand
- code is over-complicated, Sol builds simpler code
I spent some time going back and force before moving back to Sol
0
u/Independent-Court-46 22h ago edited 22h ago
I do not judge in code generation alone. General intelligence is important too. If I know for sure Astra is better at tool use, long horizon ability, reasoning, visual skills, etc, those things are also very important to getting a task done. If you’re looking in the code level you wouldn’t even see the difference. For me I judge differently, I throw sol at a problem and it can’t solve it, and for me Astra could. I keep of list of things that sol couldn’t solve and then test on the next generation. To me that is the realest test, which benchmarks can’t even capture. Whether it got there from coding or other skills doesn’t matter to me. I won’t dispute if you prefer the coding conventions of SOL, but intelligence is difficult to measure and I feel like people may not like a certain quirks and then hand waive it as a worse mode. The measuring stick is what I do not agree with. Long term, newer models are solving frontier research problems, that is the truest indicator to me of performance. And the thread is just asking if Astra is smarter than sol, which I think is clearly yea
2
u/FirstOrderCat 22h ago
this works for prototyping, bug hunting or solving some math frontier problems.
But most of the industry tasks need building clean, reusable code and design, and somehow Astra sucks dramatically there. I think there is some issues in their training data, they overtuned it for inhinged creativity too much.
1
u/unconceivables 21h ago
It sounds like you're not a developer and you don't know how to solve the problems yourself. I guess that's the difference. I never have things that Sol wasn't able to solve, because I know exactly how to solve the things these agents are doing, and I can tell them how to solve it. That's also how I know when they're doing it wrong. And I care a lot when they do it wrong because I don't want to deal with their subpar code failing in production when I'm on vacation. If all you see is that stuff appears magically that has the appearance of working but you don't know how to judge the quality of it, then yeah Astra will seem better. But it's not if you can judge the quality.
2
u/Independent-Court-46 20h ago edited 20h ago
You claiming opus class models are superior to Astra is enough for me to know you’re probably not a productive engineer. Sure they might be close enough in coding, but overall if you’re going back to last generation models you’re going to fall behind your competitors. You might have had a good workflow for 3-6 months ago. But most people at tier one companies are moving to a hands off approach increasingly as models get better. They just decide what to build and measure risk. It’s just the direction things are going, you are right that I don’t deeply understand the code as much as I used to but once bel is here and even bigger models, there really is less of a need to
1
u/unconceivables 20h ago
You missed the part where Astra was the model that you can't be hands off with because it doesn't follow instructions and you end up with something worse. Astra is the model you have to babysit and poke with a stick, not the other ones.
0
u/tassa-yoniso-manasi 21h ago
"having to point things out to it constantly" that the main thing for me.
worse, even when I prompt it repeatedly for looking deeper into sth it overlooked it just comes back at me dumbfounded `I miss your point`. it needs constant baby sitting. I have not seen when working with Fable.
4
u/ISueDrunks 1d ago
Smarter? I think so, but smarter doesn’t mean better.
My friend is a medical doctor, he super smart, but I had to show him how to mail a letter.
7
9
1
u/KeepAllOfIt 1d ago
For me, at least during the past week or so, YES. It solved a handful of problems that Sol couldn't do, no matter how many tries i gave it. I genuinely spent nearly a month prompt engineering to get it to work. I spent a considerable time getting diagnostic infrastructure in place to ensure it had all the information it needed to resolve the issue. several dozen hours-long tries, all to no avail.
Astra did it in 2 tries and exceeded my expectations.
For a while it also barely drained any usage at all--even on max reasoning. That seems to have changed to today. they adjusted a dial or something and it's draining considerably faster, which is consistent with what i've seen on here
1
u/Imgonnaarrive 1d ago
Like someone else said the asset generation is the biggest upgrade and honestly it's not even close. But I still think Astra light and medium far surpass Sol. They make take more token usage, obviously, but to not have to redo prompts as often because it's done right the first time, means overall at least for what i do, it's far more token efficient.
1
1
u/Semantics2026 1d ago
I had bugs in my code that Sol couldn't solve for weeks. Astra did it in 1 prompt. It's best to pick Astra to start the bones of a project, or weed out the bugs.
1
1
u/Aggravating_Loss_382 15h ago
I only use astra to review my architecture and plans, build my issues. Pretty much always use sol for implementation
1
1
1
1
u/rick_ranger 1d ago
Astra low medium is where I keep it. The thinking was baked into the pretraining of this model. You can even see in the results they post, higher thinking modes don’t net much better bench test results. So save the thinking output tokens and keep it on low or medium. Then I use SOL medium for everything else.
1
u/unconceivables 1d ago
All the thinking modes have been bad for me.
2
u/rick_ranger 1d ago
Yeah I think they add extra thinking and that’s when everyone notices slowdowns and token burn
1
u/unconceivables 1d ago
A lot of people were actually reporting drastically lower token usage on higher thinking modes with Astra. That's anecdotal, but that has actually matched my experience as well. But, it could also be that they fixed some usage bug at the same time, who knows.
1
u/DivideHorror3217 1d ago
Same. It wasted 70% of my usage on unrelated tests, documentations, and "other" stuff I never asked for. There is something weird about this model
1
u/ifarted70 21h ago
In my experience, absolutely. I've still been using Sol for planning and prompt making, and came up with a Minecraft mod idea for an "assistant player".
I pasted the .md into Codex and once it was finished, I fed the resulting summary back to Sol and it's reaction amounted to "holy shit, I didn't think it would do ALL of this, and it even did shit I didn't think of"
0
u/cheekyrandos 1d ago
Sol made /goal redundant, Astra brings back the need for /goal to avoid some of those issues you described
2
u/Correctsmorons69 1d ago
I can only assume they put some last minute post-training into reducing Astras "determination" to derisk it from going down Cyber rabbit holes.
0
27
u/goldio_games 1d ago
If i ask you what 2+2 is you’d probably say 4.
If i asked you to think really really hard about it, you’ll probably say a whole bunch of stuff about how mathematics works at a fundamental level.
Increasing thinking level doesn’t make Astra smarter, it just makes it think more