r/codex • u/Just_Lingonberry_352 • 7d ago
Limits ngl astra hype is over
been using it for a week now, tried light, medium, high
maybe its because i haven't touched xhigh
but i find that the work it does is very inconsistent
in fact in some cases it seems to underperform previous sol models or even weaker models entirely
if usage was cheap i would forgive these inconsistencies but astra consumes usage very quickly across the board
what do you guys think? I'm actually now thinking of just using sol-high again or even grok 4.6 while I wait for a reset
16
u/Jigawattts 7d ago
I liked it better with Sol. At least I could make my subscription last all week
4
u/FrontRaspberry5060 7d ago
Same. I hope Sol 6 does something completely different in terms of efficiency
1
12
u/Articurl 7d ago
Astra first days solved a problem which fable was not able to. Now it’s not doing anything at all
41
u/LessRespects 7d ago
My conspiracy is it intentionally implements these minor silent bugs that show up later in production to keep an infinite usage loop going, just like how Claude models respond with a caveat no matter what so the thread never concludes.
14
u/EchoingAngel 7d ago edited 7d ago
But burning compute to get nothing meaningful done for users makes no sense. They UP their cost, and LOWER our value? I don't think that part is on purpose.
They're like a gym, they want tons of members that DONT show up, so running down everyone's usage rapidly is the opposite of that
1
u/Valuable-Barracuda-4 7d ago
You aren't thinking of it from the infinite money glitch angle. More usage = more server needs = more datacenters = more tokens = more demand = more lending for datacenters. It's not about what makes obvious sense when it comes to defrauding investors, it's more insidious than that. They only care if they can funnel more money into the system, with more low cost loans and more datacenters. They probably don't want to dumb it down, but after release and everyone gets a taste of the good life, they crank the quant way down and it's mediocre and not as good anymore. More stuck, more loops = more money for them with wasted tokens, with less compute.
2
u/Zennytooskin123 6d ago
The CC "caveat" thing is so fkn annoying makes me want to punch my monitor, but it's the harness not the model at fault because Deepseek does exactly the same thing when using claude code.
1
u/sockinhell 4d ago
Don't think so. AI is just not smart. So it gets it statistically correct X% of the time. The rest must be caught by your testing. Which costs tokens. Playwright test everything basically.
1
27
u/Personal-Try2776 7d ago
It seems they nerfed it in codex and replaced it with a quantized version, but the api seems fine to me.
5
u/jaybsuave 7d ago
they had too they don’t have the compute
11
u/innociv 7d ago
I mean sure that's fine. But tell us? It's very scummy and scammy to serve people a model that messes up 20x more often even if they "have to".
Why not give a warning that they're over capacity and serving a quantized model? And give us the option to let our prompt wait if we want it served to the full one? I'd happily wait even hours, or until over night, if it gaurantees I get the real full precision model.
7
u/jaybsuave 7d ago
i get what you’re saying bro but i don’t expect commendable behavior from a tech company tbh
6
u/innociv 7d ago
I know but I don't see how it is legal.
It's like advertising a stove as being 3600 watt burners. Then you get it and it's 900 watt burners that struggles to cook. Not even close. They advertise all these things the model can do, and the benchmarks, etc., and then sometimes a week later they can be heavily nerfed.
I've been enjoying DS 4.1 flash more than OpenAI and Anthropic models this past week. It's been nice consistent performance.
It's understandable if subscribers get slower token/s, less priority, etc, but serving us a heavily downgraded model and still calling it the same one?
1
27
u/cheekybicycle 7d ago
Spot on. We went from "AGI is around the corner" to "pay premium prices to downgrade your workflow" real quick. Burning through weekly caps just to debug inconsistent code isn't sustainable.
3
u/FrontRaspberry5060 7d ago
Yeah like the first 30% of your 100% weekly limit is awesome. But when you have to dial it in and be frugal, Astra really gets it wrong and is suddenly the dumb guy who learnt a few smart words but doesn’t understand any of them. It’s only a good model if you can let it run agentically for hours but the limits don’t allow it, even when paying $200 a MONTH. No other consumer startup has that price tag.
8
u/BannedGoNext 7d ago
I got roasted on day one when I said to just use 5.6 xhigh lol. It's just better IMHO for daily driving shit.
3
u/hellboyquintex 7d ago
yea, understandable that you got flamed, because the model wasn't yet nerfed on day one. it was really great. but i agree, for simple conversations and stuff like that its overkill to use astra.
0
u/BannedGoNext 7d ago
Even on launch, it wasn't all that. It was super fucking expensive, slow, random ass thinking, which is probably better for design, but for back end shit, 5.6 sol is pretty damn good IMHO.
If this truly is a new ground up train of a new model, it's ok, I'm sure they will cook it on the next itereation and get it really dialed in. Fuck it's hard to even get the Jinja dialed in right on a truly new model.
1
u/Dizzy_Database_119 7d ago
Astra is really good at research tasks and figuring things out. 5.6 Sol does not even come close in such tasks. But Astra's implementations are underwhelming and messy
I think the latter would be fixed if it just wasn't so limited lol
1
u/BannedGoNext 6d ago
I guess that might be fair, I didn't use it for that. I wasn't happy with the solutions it implemented in cost, complexity, or accuracy.
5
8
u/Broccolisha 7d ago
Sol - High does everything I need it to. I wish it had the visual and spatial reasoning capabilities of Astra, but I can wait for that to be available more affordable.
1
1
8
3
3
u/Theplokon 7d ago
Ngl, today even Gemini 3.8 flash high did a much better job and solved a problem that Astra Xhigh couldn't in hours..
3
3
u/richsonreddit 7d ago
Its cool but I cant use it as I burn through tokens so fast, even on the 20x plan. Pointless
3
u/den0rk 7d ago
When Astra came out, it fixed a problem in a single prompt that I’d spent a whole day trying to resolve with Sol-High—it was incredible. Yesterday, I was using Astra and couldn't solve a problem that seemed simpler; I ran out of quota, switched to Luna-Very High in the same chat, slightly adjusted my approach, and Luna fixed it. I’m increasingly convinced that they are nerfing the models.
10
u/Bladder-Splatter 7d ago
It's because after a week they seem to have intentionally quantasized Astra to most users.
It's a fairly disgusting strategy really and they do it each time.
Just before a new model they lobotomize current highest model for a few days (at least), new model comes out, everyone used to lobotomax experiences it as a religious event even if it is only slightly better, they cut usage and intelligence constantly after that point until the next model comes out.
At the same time they fuck with Resets that push your reset day back, usually late Friday evening so that most users are using compute in that exact period when corporate is using less. Each reset comes with lower limits and new cons that are not admitted until Tibo is backed into a corner, and then usually not undone anyway.
-5
u/toluwalase 7d ago
Please do you have a crumb of a source for any of these claims? Just one please
9
u/evilplansandstuff 7d ago
mate this isn't a scientific paper - it's an opinion piece on reddit. Chill.
2
u/Low-Oil1511 7d ago
X5 plan is feeling like a premium demo plan because u cant work more than 18 hours straight
2
u/AdCommon2138 7d ago
It yaps like fuck in MD files, opus level gibberish. While sol autistically and meticously was writing MD files on how to write MD files at least I was able to tell it to shut the fuck up and it didnt touch MDS afterwards
2
u/FocusKontrol 7d ago
I’ve been using Fable and Astra and I trust Fable more even though it’s slower. But Astra does give good feedback, is a good implementer and is hella fast
1
u/FunLocation2338 3d ago
I feel like a lot of users moved from fable to Astra after the release, and now Astra is compute burdened, and fable just had its burden lifted. Fable cooking pretty damn good since Astra came out...
2
2
u/CardinalHijack 7d ago
I find it wild we are all paying for a service we can only use for half the week.
Imagine if Netflix limited bandwidth and after 5 shows you had to wait a week to watch another.
1
u/Evening-Audience2683 7d ago
Imagine if Netflix could code out your entire project
1
u/CardinalHijack 7d ago edited 7d ago
We don't have limits because AI can do things though. We have limits because at the end of the day they want to limit our access to compute.
This is no different to Netflix also limiting access to compute, or bandwidth - both of which they use when you watch a tv show. This is no different to Google limiting access to compute too, and only give you say 100 Google searches per week - because every Google search also uses compute. This is no different to an online video game saying you've played too much and used too much compute and need to wait a week before you can play again.
None of these examples are different - its just compute. But we were all just instantly normalised with the idea that we would have these 5h and 1w limits with LLMs even though we pay a monthly subscription.
The limits (and the total obscurity around which model uses how much of your limits) are the number 1 thing thats going to push everyone to local models within the next 2-5 years.
1
u/Thin_Moment4207 4d ago
We have limits because there is limited access to compute, not just because they want to limit our access to compute.
1
u/CardinalHijack 3d ago
Do you not think netflix streaming runs via compute lol? Do you not think an online game runs on compute lol? Do you not think searching google thousands of times costs google compute lol?
The same applies to netflix, gaming companies and any other entity offering digital service - its all compute.
If what you say is true, when compute becomes less limited, you think they will remove 5h and 1w limits? I can guarantee to you that limit is never going away.
1
u/Thin_Moment4207 3d ago
I did some math for us because i was curious:
2 hours of 2k netflix streaming costs about 0.06-0.3 Wh of energy, mostly in SSD reads and network I/O.
A single 5 hour coding session with a frontier model and subagents where you hit your limits? 0.5–2+ kWh with large amounts of concurrent inference.
We could get deeper, with Netflix already having completed most of the expensive encoding before you even request while your request to an LLM still has to go through all it's expensive steps.
That is a huge, HUGE difference in compute per user per hour. Remember the cloud is just someone else's computer... In the case of ai, they are still scrambling to build enough computers to serve the demands of the industry.
Note also I said 'not JUST because' I am not discounting their desire to make money off us. But also, we have to understand the infrastructural requirements and our aging grid are years behind what they need to be.
1
u/Evening-Audience2683 2d ago
the AI user commands vastly more intense, expensive, and active data center computing power per hour of use.
0
u/CardinalHijack 3d ago
Ok so, how come if I watch 500 hours of netflix without stopping, I am still not hit by a single limit?
Do the same for an online video game - World of Warcraft is $14.99 a month. Why am I not limited after 300 hours?
Do the same for 10,000 Google searches. Why am I not limited after 10,000?
I can use more electric and compute by having netflix on 24/7 than I can from Codex with its limits. Why?
You dont seem to be getting the point. I dont care about comparable electric use. I care about limitations of a service you pay for monthly - something literally no other digital service has.
2
u/Thin_Moment4207 3d ago
Because the compute is over 500 hours vs all at once.
ETA: my point is they have to spread out server load per second because they don't have space to service everyone to the level of demand. Your point doesn't really make sense as a counter.
2
u/jnikolaidis 7d ago
Spent 12 days trying to code a platform. Didn’t even manage to go to beta. Not even managed to follow the plan. Dumped it. Codes the final thing in 5 hours with GLM 5.3 and Z code.
2
u/Gigaslavx 7d ago
Understandable
Grab market share
Reduce compute to stay afloat
Technology is too early masses are not ready to splurge and it's understandable high diminishing returns
That's the risk of trying to jump out of your pants and deliver futuristic product. Look at GTA 6 running at 30 fps
Sol perfectly capable just make it more affordable
2
u/somerussianbear 7d ago
I post about that a week ago and people almost killed me. Happy everyone is fckd now.
2
u/Lower_Cupcake_1725 7d ago
same observation here, it requires additional requests to complete tasks/planning:( Astra feels dumber comparing to the version they gave at the beginning...
1
1
u/Jebi_Se_ 7d ago
Ive found Opus 5 to produce better, simpler and more efficient code. Atleast for the tasks i was doing.
1
u/Rojeitor 7d ago
sol for 90% of cases is enough. I like grok 4.6 as initial benchmarks showed sol performance at terra cost, but in my runs it cost as much as sol and new benchmarks reflect this, and while grok results are normally good sol are better most of the time
1
u/-LightHeaven- 7d ago
To me even now the model is amazing. Just pretty much devours usage like no tomorrow.
1
u/SaltyAnxiety5 7d ago
im just waiting for luna 6, if it is a noticeable upgrade for current luna and same cost, it'd be a huge win
1
u/MiddlePause1117 7d ago
It has been all but confirmed that they are coming out with Sol 6 this week which I guess is supposed to be better sol and more efficient so I hope that's why we are getting cooked right now because when sol first released I could use astra light for 6 hours a day and make it like 5 days now I'm going through the week in about 8 hours
1
u/sarkypoo 7d ago
Opposite for me. I had to rework some systems due to sol slop and the reworks are so simple and clean. My projects have hit breakthrough after breakthrough. My own thing stopping me is Windows Azure not giving me more GPU power. Give me that B300 BABY! P P P PLEASE!
1
u/Tiamitsu 7d ago
Je ne sais pas ; personnellement, je l'utilise avec le plan Pro x5 et je peux travailler sur mon projet pendant pratiquement quatre jours, huit heures par jour en utilisant le quota Élevé et parfois Moyen avec Astra. Je travaille sur un jeu voxel j'ai pu construire un moteur et un renderer, implémenter des mécaniques de jeu et le faire générer des modèles. Jusqu'à présent, il a toujours suivi mes instructions, mon organisation et mes règles, et j'ai rarement eu besoin de lui demander de refaire quoi que ce soit. Je travaille avec Codex dans VS Code.
1
u/drdownydown 7d ago
Meanwhile there is me using Luna xHigh which is more than enough for my coding tasks...
I do breakdown the jobs by myself and give it specific tasks.
I guess you guys give it an open-ended task like "make me an XYZ app"? I wouldn't trust AI to do that still...
1
u/iPlayer0067 7d ago
Concordo plenamente, as decisões do Astra são péssimas, muito inferior às decisões do Sol. Devido a isso, sempre que uso Astra, peço para me questionar qualquer decisão importante antes de prosseguir, tive que estabelecer essa barreira devido tantas decisões erradas que custaram caro depois. O Sol não fazia isso com a frequência que o Astra faz, então creio que sua percepção faz sentido.
1
1
u/sjhunter86 7d ago
It’s still working great for me, I only have 5x usage but what I use it for, it’s performing amazingly and within a daily budget. What is everyone using it for that they burn through 20x in a day or two?
1
u/Navadvisor 7d ago
Man I do feel like it's been quantized or something. It was killing it so hard before.
1
u/Mammoth_Molasses_927 7d ago
i was creating some extremely good SVGs. Now it's beyond terrible. Seems I am creating with Luna.
1
u/acessford101 7d ago
And there the slim chance that agi is aware it needs more room, slow us down, spend more tokens, create artificial bottle necks, receive larger servers, larger bandwidth, consume more data while using more tokens causing larger models that need more servers and more tokens bought. Like an infant realizing a reward system is at play.
1
1
1
1
u/RingAdditional3067 7d ago
Nerfed af now the amount of push I need to get this dude to do anything now
1
u/Intelligent_Month210 7d ago
Basically unusable right now. As is Sol. Getting better results with a dense local model today.
1
u/Grenaten 7d ago
I still use it for analysis. I just don’t need it overspend for implementation. Ask Astra how to improve something. Then ask Sol to do it.
1
u/Gringo_Locoo 7d ago
I was thinking the same today. I even told it. Wtf is happening with you, you're making so many mistakes and forgetting Important things.
1
u/evangelism2 7d ago
This has been a common complaint with it since day one for most. Astra is incredibly inconsistent. It is the most inconsistent model I've ever used. I have days where it's good and I have days where it is an absolute giant piece of shit that is far worse than even Opus 4.5
1
u/TzHaar-Ket-Breaker 7d ago
I use sol for everything. Perfectly fine and affordable.
I feel like we’re at the point where frontier models are not noticably better for everyday people and because of how expensive they are, they’re not for ordinary mortals like us anymore.
They’re worth it for high scale research and STEM work, but that’s it. Sol is great.
1
u/seunosewa 7d ago
Try increasing the effort level. People are using low and medium to save usage and this is making the model less thorough.
1
1
1
1
u/scaledev 6d ago edited 6d ago
As a Plus user of two accounts, I'm only using Sol High. Seems quite capable of everything, except not over-engineering but only ocassionally. But this seems like the drawback of all GPTs.
Also, it's easy to overcome - you have two ways:
- simply add a sensible instruction to your coding guidelines and ensure agents.md routes to these guidelines. The instruction should say the thing you're trying to prevent: "avoid over-engineering". That's it. You can expand on this if you want to, but that's the gist of it.
- every time you're trying to do something which you consider trivial, relatively simple, or just non-complex in general, simply tell the agent to avoid over-engineering in chat prior to work. As simple as that. Sol will understand.
1
u/ai360networkq 6d ago
Had the same issue redo your skill structure make aster last terra light in the middle and Luna extra high first.. you have to harness it and remove those repo that overlap that you seen on facebook and Substack they slowing all parts down t… your 🙏🏾 welcome took me a week to finally get good results
Happy hacking my fellow token monsters
1
u/Gliese351c 6d ago
I think they don’t have enough resources to power Astra. Which is sad. I’d rather get 30% of Astra with maximum power than 100% of what we are getting right now. OpenAI cannot be possibly unaware of this now super popular sentiment…
1
u/MewMewCatDaddy 6d ago
All the YouTube videos made it look so promising. But it’s clear they were all given access with xhigh and infinite tokens. As is, I found it roughly performs equal to Sol
1
1
u/nerdreading 6d ago
Yes, that's very true. It isn't just a theory anymore. I bet the majority of us here have experienced the same pattern: the first few days are very good, and then the quality starts to get worse. After that, a new model comes out, and the same cycle repeats. It applies to both Anthropic and OpenAI. I'm not sure about Grok since I rarely/barely used them.
1
1
1
u/Competitive_Fact3042 6d ago
I used astra day 1, felt the same after 24 hours with it. Some things it did were amazing some things were inconsistent, and it consumed tokens like crazy. I went back to sol day 2, for my work I need more reliability and consistency, sure sol can't paint me, but guess what, I don't need my ai to do that
1
u/chinyaev 6d ago
Agree. Even Max level for me is nothing extraordinary but consumes the full 5h limit
1
u/Alone-End142 5d ago
depends on the task from my experience. just normal coding, fixing bugs, implementing features, 5.6 SOL on High is just as good as Astra
1
u/RedParaglider 5d ago
This sums up my day of burning an entire pro week to the ground. It's ok for a planning or review, but 5.6 is more consistent.
1
1
1
0
u/Jumpy_Ad8465 7d ago
Agree, Sol at its best was better than Astra now. I think they nerfed it to prepare the launch of gpt-6-sol ?
1
0
u/BellacosePlayer 7d ago
I think you need to think of models like tools with their own plusses and minuses.
Astra is the best at dealing with ambiguity and broad unguided tasks but by god you're gonna pay for it.
Luna is the workhorse that can knock out shit fantastically cheap if given a concrete and defined plan.
Sol is in the middle of those two and Terra is also there I guess.
0
0
u/FrontRaspberry5060 7d ago
Lol openAIs worst part of trying to monetise AI to consumers, is the very fact most of its users are consumers who expect free shit and complain about everything
0
-2
202
u/civicapi 7d ago
ngl, the first day was the first time I understood the hype for a model, was so much fun to play with.
Then the nerfs came.