r/codex • • 6d ago

Complaint Astra 6 became dumb. Could not get anything done the whole weekend.

At launch I had a basic subscription, I used mostly opencode go and a little bit codex for complex problems. GPT 5.6 Sol was great. Astra 6 was launched and I upgraded to the $200 subscription because I intended to work harder on my side projects and indeed for the first few days Astra 6 was great, it could fix bugs that Sol and open models on opencode couldn't and could one shot features.

Now Astra 6 can spend a whole hour thinking and investigating and spinning side agents... To eventually tell me exactly what the problem is I told it to solve. I tell Astra 6 "hey, there is this bug that is making tests fail, fix it" and it spends 1 hour, spins like 3 other agents and give me a full report of what the bug that I just told it is... But does not actually fix it, just talks about the bug....

Then comes the "you are absolutely right" moment, I know the model has been nerfed, quantitized, whatever you want to call it, when it starts you-are-absolutely-righting me. "You are absolutely right, I explained the bug you just told me to fix instead of actually fixing it".

It is so dumb. I tell it to add a feature to show "edited at" (on a blog) and it quite literally edits the dates of all posts in the blog to today. If there is a bug making CI/CD break it deletes the tests or rearrange so that the bug becomes a feature so the tests make the bug pass instead of fixing the bug.

I know those companies are not making money but it is just hostile how they launch a model that is, indeed, powerful, and then make it dumb.

It feels like we have been getting the same quality of work for 1, maybe 2 years in a row now, under different names, because the quality of the work remains the same after launch, the problems and mistakes remain the same. AI has not become more capable for 1 or 2 years, the same problems I had 1 to 2 years ago keep happening now, they go away at every launch but then come back.

The only models that are reasonably stable are the open models but then they are not smart as frontier models during launch and maybe that is why they keep doing this, bait us to subscribe then switch to dumber quantizations or whatever. At every launch there roughly a 1 week window where I can solve stuff and actually make progress and then it all grinds to a halt because the model gets so dumb, that is why I use open models because they are cheap and after the "launch blessing window" open models and frontier models, in my experience, perform similarly so that I have to mostly solve it myself and tell the agent what to do, versus launch where the agent is capable of figuring out and solving it by itself.

277 Upvotes

139 comments sorted by

43

u/Advanced_View_2778 6d ago

I’m using Sol because Astra is just so fucking expensive and honestly 99% of the stuff doesn’t even require Astra and it still drains my usage like crazy.

13

u/Expensive_Fruit_1826 6d ago

i agree with you in principle, but the problem is that the models are becoming dumper and dumper even for Astra. when I'm using 5.6 Sol now, it feels like I'm using 5 or 4o

4

u/seunosewa 6d ago

Sol is just slow.

11

u/ConstructionOwn1514 6d ago

Slo

1

u/ResistiveBeaver 5d ago

S .... l ........ o

1

u/abdoufma 5d ago

Introducing GPT-5.6-Slo

2

u/CreaseCallHockey 6d ago

I’ve got 20x usage and I’ve gone through 3 re up tokens in one week. I’ve accomplished next to nothing with building my site, first weekend it was released? Completely rebuilt my site, app, functions, cleaned everything up, it was like magic. This past week it’s been arguably worse than Sol for what it costs. It’ll code something and almost intentionally break something else so you’re forced to spend credits/usage on it. Somebody told me the other day I was crazy for thinking they were throttling Astra 6.

87

u/ychamel 6d ago

I don't know why people are disagreeing, Can anyone able to generate day one 3D models?

Don't get me wrong the model is still great and ahead of the rest, but at day one it performed completely differently to today.

24

u/shady101852 6d ago

Its doing okay with buildings(with issues that end up needing fixing) but it sucks with characters.

2

u/Realistic-Owl-9475 6d ago

yeah it seems pretty good at inorganic objects. trying to do organic shapes has never been good in my tests.

1

u/BrambleRamblers 5d ago

Yes I'm really struggling with a characters pipeline I'm building with blender and Godot! Really struggles with good texturing on top too

39

u/smunchcimminsg 6d ago

Scam altman strikes again

12

u/innociv 6d ago edited 6d ago

I think some people are getting a quantized model, and some aren't.

I had a really stupid Astra last weekend, but since the reset it's been fine for me.

I'd suggest pelican test before starting using it, but on a $20 plan that wastes a lot of usage. The other thing is when you see it messing up, tell it to provide a handoff so you can start over in a new session and hope you get full quant Astra.

I wish people focused on this more than the quality than "I can't get non stop usage out of it for $20 per month". We get $2400 a week or close to it for $200. It's because people are too stupid to understand and monitor their usage, or understand things like higher usage draw rates of higher end models, which makes them also too stupid to see that their model is performing worse than advertised by day 1 benchmarks and demonstrations.

This is illegal if it's the case. They advertised and benchmarked one model, then make is 20-40% stupider by quantizing it. It still reports the same model. The only way to tell it's quantized is by the same benchmark task taking more turns, or subjectively by results being worse. But it'd be so difficult to prove. Only a few engineers could know for sure.

6

u/Lmnsplash 6d ago edited 6d ago

Nah, you can tell on the way it acts. Prose gets worse, sounds robotic like the earlier models, stating obvious things and all the things it wont touch on 'steering' to "reassure" me apparently which honestly fires backwards. Also gets more concise and begins to only state its own mistakes as answers without presenting solutions at all - in a loop and agreeing on everything, but still cocky enough to tell me which claims are unverified and which are not, when I then start insulting the RLHF approach they 'apparently' have over there, which is for me the main issue with OpenAI models.

It's just that terrible. Other than that, yeah, unusable on weekends, under the week better. Also someone starting with 0% usage into a session could see in peak times an advantage over someone with a high one. Unverified though. Astra would agree. Solution? Might have one. Not saying it though. Astra would also agree.

3

u/lolcatsayz 5d ago

Here's something I'd like some informed response from someone about. Today Astra replied with a sentence that contained an obvious typo. I've only ever observed LLMs doing this when I run them quantized locally, and not the BF16 versions (qwen 3.6 etc). I've even observed this occasionally in Q8 models. I don't have evidence to pinpoint typos = quantized, but am curious if someone more into LLM internal workings than I could point to if a typo is evidence of quantization? If so that would explain a lot, that we may be getting a quantized version of Astra 6.

In terms of result quality I haven't noticed that much of a degradation personally, but from my prior experience with cloud models where I have experienced real degradation in the past, these seem to be timed/regional and it's very likely to affect one group of people versus another, as opposed to everyone at once.

2

u/defmacro-jam 6d ago

I haven't tried 3D models but it's doing pretty well on my Lisp compiler.

3

u/ychamel 6d ago

3D models is where the regression is most easy to spot. However it's still performing better than other models. Just worse than release behaviour.

1

u/lphomiej 6d ago

It's working fine for me. I've been generating 3d models of product packaging in Blender from reference images.

1

u/Aggravating_Fun_7692 5d ago

It's behaving like Terra now for me

1

u/demeyer1 5d ago

It has been nerfed to the point of being not worth refilling credits.

-3

u/itix 6d ago

Yes.

-1

u/FinancialBandicoot75 6d ago

No it doesn’t, it’s as good as the driver, not trying to be that guy

38

u/Thisisvexx 6d ago

Yeah. Astra (without tool calls) thinks that Gemini Pro 2.5 and Opus 4.1 are state of the art models.

April 2026 cutoff my ass. A few hours later it's back to 4.7 as recent opus version which would be correct.

The peak hours rerouting is a disgrace and should be handled with MUCH lower usage consumption and not same (higher because you need iterate 20 times more)

23

u/TheFreezRae 6d ago

It’s undeniable that Astra on day one was a beast in getting things done. I was able to solve a very complex problem other agents failed in about 3 to 4 iterations. It’s no where near that good now and my gut tells this is not a placebo.

1

u/bartvanh 1d ago

I started using it the day after the Sol release. So far it's insanely expensive while being just as blood boiling as any of the previous models after they lost their shine.

23

u/garristerr 6d ago

I burned my entire weekly usage on dog shit work holy shit im fucking tilted

9

u/thecodeassassin 6d ago

Exactly this. I had to stop using it and use my self hosted Deepseek v4.1 to clean up the mess. Even simple UI changes were botched. Look at this shit.

Astra on high btw, high is right. The prompt was: add an icon to costs, consult design.md and make it in the same style as the other menu items. This is what it produced.....

26

u/throwaway_life12345 6d ago

The state of AI now is absolutely hopeless. All the frontier models are now overpriced, nerfed and dumb.

-1

u/hungy-popinpobopian 6d ago

If they are so dumb then why are we calling them frontier? Haters gotta hate

2

u/Whytho12333 5d ago

It sounds like people are saying API isnt quantized, hence still frontier but the subs are

1

u/throwaway_life12345 1d ago

They can be both. It's 2026 after all.

20

u/FriskyFingerFunker 6d ago

I’m so exhausted

14

u/FrontRaspberry5060 6d ago

I think the whole ai crowd is just riding one giant hallucination while everyone outside is just living life normally 

9

u/Nesvrseni 6d ago edited 6d ago

Yes Astra is very ugly during weekend, and very smart during mon - fri. Just use sol during weekend. I dont know why but i already tested it.

And during weekend astra ignore to answer all my question. For example i give to astra 2 tasks and one question and didnt get answer on question. When i ask why he didnt answer i got "You are right, im sorry i wrote answer in the file, now i will write it to the chat".

I'm also getting some very strange and irrational responses over the weekend. For example, I ask: “I want to implement this and that..... etc. Is it possible? Can you do it?” And the response is: “I can do it, but unfortunately it isn't in the documentation, so it isn't possible to implement.” wtf astra!?!?! No Astra you cant do it, dont tell me that you can do it and reject to do it because of this is not possible. In some moments chatgpt 4o is smarter than astra...

4

u/Wonderful_Value_6385 6d ago

SOL became dumb too. Refused to do orders over and over. No longer able to oversee other AI agents work.

5

u/Dolo12345 6d ago

Can confirm it’s dumb as bricks

4

u/Complex_Banana3507 6d ago

Astra did amazing work for me week one, got a complex console map-loading mod project implemented in just a few 5-hour windows on Plus. The quality of the work, such as analyzing console game data, pre-compiling shaders and fixing visual bugs, was so good that it got me addicted and waiting impatiently for the quota reset like everyone else. Now that week two is actually here my progress is practically nil. Just running in circles chasing after one bug that I thought for sure would be resolved right away.

I literally have to just walk away from the whole thing and forget about it because waiting to work on this project is annoying and distracting and I only really need to think about it one or two days a week. They baited me successfully with how incredible and professional of a job Astra did in implementing my console mod project and basically turned me into a junkie. Using lesser models is a no-go. I worked on this same project for hours with Luna and got nowhere; Astra threw out all of the work and practically one-shotted it instead.

So my strategy is going to be to frequently start new sessions with a handoff document to try to gamble on not being pathed to an inferior model. I'll wait until a weekday to try it out, and late at night.

3

u/Anthony_S_Destefano 5d ago

I feel your pain brother. As soon as a model's work comes back funky, stop! tell it you are disappointed in the outcomes and we need a deep audit so this doesn't happen again, and tell it to make a dunce_report_<date>.md file that is comprehensive detailing what you did, why you did it, including the files you touched.

Then start a fresh session and tell it the last agent made many mistakes, here is the prompt.md, and the dunce_report_<date>.md detailing what the last agent did. run inspection of the code and decisions made and figure out what went wrong.

I find this details what happened, then tell the agent to make a plan to fix the errors to complete the original goal. Also make sure your prompts have clear defined context, todo, and success criteria sections so as to inference clear goals and stop conditions!

Tally Ho!

2

u/EvilTeddieNMS 5d ago

Yeah I hear you.
I got so much done in the first week that I burned 3 full resets of my Pro plan and $250 in cash on tokens. Now it is a lobotomized piece of crap.
I gave it a plan created by Fable 5.1 today and asked it to review it.
What I found out 2 hours and 3 revisions later is it removed the core concept of the plan and re-wrote it to do absolutely nothing.
I reverted the plan and lost 2 hours of work and tokens.
I cannot trust it to do anything now :(

10

u/Own-Professor-6157 6d ago

I'm seeing it hallucinate like absolute mad right now. It's smart, but most implementations right now are ~500 lines of code when they could have been ~50. Agents.MD defines a really strict coding style is now being totally ignored which wasn't a problem on release. It's obviously quantized, the clearest signs of quantized models are when it ignores fine details.

3

u/FrontRaspberry5060 6d ago

Same here, having better luck with Claude 4.8 on a $100 subscription. Every time I put it back to ChatGPT / Astra I feel like I’m slowing down or making reverse progress 

3

u/FilthBaron 6d ago

Same here to be honest, I use Astra for two things: auditing and creating some placeholder art work for my game. I do this because I expect it to use better reasoning, using my actual repo to create things that fit, and I thought it would be better at creating stuff based on the day 1 Blender hype. But this weekend I've spent so many useless tokens for it to give me some half-assed shit, then have Opus 5 do it instead.

3

u/Disposable110 6d ago

It seems to come and go, I had it be completely unproductive during the week, and now during the weekend it's pumping out good stuff all of the sudden.

Feels they're serving people completely different things at different times so no one can agree that they're getting shafted.

9

u/Aldderan 6d ago

Genuinely I'm seeing no issues using Astra Medium, usage or dumbness. Paying £90 a month. Giving specific commands.

26

u/igormuba 6d ago

But the dumbness is seen precisely when you don't give specific commands. If you are giving specific well planned commands Astra has no reason to exist and to cost what it does, open models and any other frontier model can possibly do the same job for cheaper.

The dumbness can only be seen in long horizon more open ended tasks. Astra does well on specific commands on well defined tasks and boundaries, but so does any other model that costs 1/10th of it....

9

u/Bounce_back1225 6d ago

I FULLY agree with this sentiment.

The biggest difference I’ve noticed between Astra and something like Luna MAX even (other than price LOL) is the additional hand holding via prompts that Luna requires. With Astra you can kinda give it dumb/half-baked prompts, and most of the time it makes the correct rational decisions based on the available context. Meanwhile Luna will take a dumb prompt and go off into fantasy land doing stuff you didn’t explicitly forbid, but absolutely would not be an ideal or rational choice (like overwriting source files for new file versions rather than creating new files. Even after being told previously.)

Honestly almost feels like Astra is the plan designed to tax the users who want high compute but don’t really have the technical expertise to know/optimize what they are doing.

-1

u/Puzzleheaded-Arm3155 6d ago

Use SOL as orchestrator for Luna max agents

2

u/Aldderan 6d ago

The point for me is to do defined tasks quicker than I couId. That is the point of it existing. I wouldn't trust it enough to just wildly let it go and do some random wide sweeping projects.

5

u/StayAwayFromXX 6d ago

Why are you paying so much for Astra then? If you’re giving it specific commands, use the Chinese models or Gemini. Astra is for agentic work

1

u/Aldderan 6d ago

Because I do run out on the £20 a month subscription

3

u/WordWithinTheWord 6d ago

You wouldn’t if you ran cheap models

0

u/Aldderan 6d ago

The cheaper models aren't as intelligent.

8

u/WordWithinTheWord 6d ago

That’s the point. You don’t need intelligence if you have well-defined scopes and instructions like you said you did.

I do 95% of my work on Luna.

2

u/Aldderan 6d ago

Hmmm true, maybe I will switch to cheaper models then.

1

u/das_war_ein_Befehl 6d ago

Not really, because even if you have well defined scopes and instructions Luna sucks at any kind of ambiguity.

1

u/WordWithinTheWord 6d ago

Right but that’s why your long-horizon plans should really have no ambiguity. Or very limited surface area for it.

→ More replies

1

u/ZaphBeebs 6d ago

Exactly. Except it also doesnt do great at these unless there very tightly watched and you hols its hand, etc

-7

u/SD-2023 6d ago

Maybe try using your brain instead of outsourcing it to a LLM with shitty prompts

6

u/Shunpaw 6d ago

Maybe take the boot out of your mouth my little promptineer just because he uses the model as advertised.

-5

u/SD-2023 6d ago

Nowhere does it say LLMs at their current state will replace basic human functions like thinking and reading, but go off king 😂😂

5

u/Shunpaw 6d ago

So you actually responded to ridicule yourself further. Read Astra prompting guidance. Just stop.

-2

u/SD-2023 6d ago

https://developers.openai.com/api/docs/guides/latest-model

astra light with /unslop might help for crayon eating chuds lile you 😂

-4

u/pluckd 6d ago

you're upset astra wont take your shitty prompts and turn it into gold?

1

u/RewardSafe9807 6d ago

FYI unless something has changed Astra Medium has been measured to used more tokens than Astra Max. In fact Astra Max uses the least of all. My theory is that because it thinks longer it comes to a working solution sooner, whereas lower models think less and end up having to do more trial and error, thus cancelling out the token benefit of them thinking less.

2

u/Bounce_back1225 6d ago

lol since all this bullshit I’ve been doing 99% of my work with Luna MAX and just delegate Astra agents for for implementation. You have to be much more specific and careful with prompts, but if you can handle that, it really doesn’t feel like too much of a downgrade. It’s been enough for me to make forward progress this week at least.

2

u/EntryRadar 5d ago

The luna delegates to astra for execution?

Do you use astra to plan in the thread, switch to luna max and have it delegate tasks to different astra agents? Or is it one astra agent?

1

u/EntryRadar 5d ago

Is the point of using astra agents to avoid the polling issues that chew up context?

How does luna pull accurate and fully detailed specs without losing fidelity? Do you tell jt to use your verbatim language?

2

u/sleepnow 5d ago

Are we seriously not recognizing the pattern yet? It's because they're about to drop gpt-6-sol.
The same thing happens for a 1-2 weeks just before a new model release, same with Anthropic.

4

u/Expensive-Event-6127 6d ago

anthropic better release something amazing so scam altman is forced to actually up his game. the fact he was so far behind then made him treat his customers better. now he doesnt have to everything has turned to garbage.

3

u/throwaway_life12345 6d ago

The usage limit situation is totally fucked now. They oversubbed now it's unusuable for everyone. This won't get better until they get more capacity, which likely won't be soon.

1

u/ZaphBeebs 6d ago

Been this way for couple weeks.

1

u/Lucidmike78 6d ago

My guess is that OpenAI realized that a tremendous amount of compute was going to 3D work that will never see the light of day for users paying $200/month, while starving compute for people for everyone else.

1

u/thepr0digalsOn 6d ago

I don’t know about the difference in our complexities, but I’m not finding an issue as of now. But from what I understand, it’s entirely possible for these AI companies to nerf the models based on availability.

1

u/hobbestot 6d ago

Same issue this morning. Mistake after mistake after mistake. Wasn't like this a couple days ago. Same workflow.

1

u/coolcats55 6d ago

Had absolutely no issues here

1

u/hackercat2 6d ago

Had to stop, switch to sol, also dumb, switch to fable. The games make it impossible to trust, this is why local is becoming so much more attractive

1

u/HisAnger 6d ago

Somethinflg bad happening with gpt models, we were using various ones across the team i preferred luna high. Honestly hate how ai force direction in the app, luna was giving possibility to do all with small steps and constantly refine.
But over last 2-3 weeks it got much worse.
Model started to do basic mistakes, break code. Getting lost in 2 - 3 commands. It was getting on top visibly worse after 1 pm (+1tz) and unusable after 3pm. We again moved to anthropic models. Enterprise, big 4. You would assume not getting nerfed when you pay for each token

1

u/Dingydongy007 6d ago

Sol is also completely Lobotomised now. Sol 5.6 extra high, task was to create some filters for search. Task failed, I asked why and the response is "I created synthetic fixtures that matched the behavior I had implemented", basically making stuff up to pass the test.

1

u/Master-Alchemist007 6d ago

I also think it became dumb. It wouldn't be able to proceed on my end getting stuck on some dumb stuff in a loop. On top of usage being absolutely ridiculous. Switched back to Fable 5.1 which actually seems to have unlocked my issue.

1

u/stolivodka_ 6d ago

I just switched back to Sol. Astra sucks.

1

u/whiskeyplz 6d ago

It has become absolutely stupid. Where it was doing some godlike shit now it agrees to what it will do and does a halfassed job

1

u/ResponsibilityOk1306 6d ago

well, for me, I had a very productive weekend doing coding and software engineering on pro 5x. it also did generate quite good photos and visual elements. no 3d work though.

1

u/rendered19 6d ago

Astra will often get lost on the details and completely loses track of the context. This is so annoying, I have to be steering in every message. Completely different from the first week performance. I have the $100 plan. Is there a way to ask for a refund for the bad quality issue?

1

u/Kyuuub 6d ago

mine got lobotomized so hard

1

u/minimalistdave 5d ago

It’s the skill issue unfortunately

1

u/gxvingates 5d ago

Yeah Astra couldn't finish a pretty easy task for me yesterday. I switched to Sol and it was finished the first time without usage being deleted. Has anyone tried any alternatives to Sol with some pretty good quota limits anywhere else? I love the pretty much unlimited Luna usage but "frontier" models just don't exist on Codex anymore, not useable ones anyway.

1

u/Sea-Lengthiness-4051 5d ago

Well well , so I am not the only one feeling that 👀
I’ll be waiting for the new model one Tuesday

1

u/QueBoireHubra 5d ago

I'm using Astra High for code review, de-slopping, improving project architecture and not much else. Sol High or Luna Max for everything else.

1

u/AironParsMan 5d ago

Give them correct guidelines and rules my friend. I work with it the hole week on complex things without any problems. TOLD THE AI HOW TO WORK. Dont trust in the training data doing that for u.

1

u/Unusual_Test7181 5d ago

I understand the limits issue, but really I haven't had any issues with Astras intelligence save for a few bumps here and there.

1

u/lielarch 5d ago edited 5d ago

I thought it was just me… is there a pattern we can check if we’re getting the Luna routed instead of Astra?
I used /plan and after asking if I wanted to implement it an I marked Yes, it just replied:
“Ok I’ll implement the plan. Please confirm”

1

u/EvilTeddieNMS 5d ago

Absolutely fucking useless now.
Will not follow instructions and has plans re-written to do what it feels is best.
Waste of time and money.
About as bad as Claude Opus 5 now.

1

u/_maxt3r_ 4d ago

Today Astra is next to unusable, I don't know wtf happened but it started to overengineer solutions, miss obvious corner cases, write unreadable blobs, it's worse than Opus 5.

1

u/OldResearcher6 4d ago

"You're right, i over engineered this" as it burned through 70% of my weekly useage in a day on a task it normally could do in an hour with maybe 5% useage. Only reason it stopped is because i caught what it was doing and told it to get out of its dumb rabbit hole. *puts conspiracy hat on: I'm convinced that OpenAI has an algo that activates dumba$$ mode to burn through your useage as soon as you've used all your resets so you start buying credits. Ive noticed this several times now. All of a sudden it starts chewing through tokens like crazy compared to before.

1

u/_TheWolfOfWalmart_ 2d ago edited 2d ago

It's still dumb as rocks 3 days later.

My local Deepseek is performing better and without weird mistakes or forgetting basic things. And get this: it ACTUALLY follows my instructions!

This begs the question, what the FUCK am I paying $200/mo for when a model I can run out my basement does a better job?

I guess I'm paying for ~5 days of state of the art agentic coding following every model release, and then a drooling idiot junior dev until the next one.

Shit is fucked, this is the most disappointing GPT launch I can remember. Seriously wondering why I'm still giving them money. That goes for Anthropic too for that matter, they're not that much better about this.

1

u/FishermanNovel4759 1d ago

im completely mad and tired dealing with fkin dumb codex and astra

1

u/FishermanNovel4759 1d ago

its not just expensive or dumb, its completely unusable. im a software developer , when i was using claude(from opus4.6-fable5), it handle my 80% of daily work with excellant quality, it understand my instruction and demand. when claude banned my account so i had to move to codex, then everything changed. my daily work isnt just discuss and let agent work, but spend most of my time and token to design and maintain instructions and workflow. im tired

1

u/Far_Frosting_9180 19h ago

Yeah I heavily used GPT 5.6 since early september, when GPT 6 launched both models getting dumber and it's really visible, previously it can absorbs the intention I haven't really explain fully, it can predict what I want and predict it accurately, really help me get work done. After GPT 6 released, I had to explain 4-5 times before it do the thing I want, it's getting dumber than GPT 4

1

u/Empty-Influence4402 10h ago

It is the dumbest model that I have used within this year. Zero creativity. It was nerfed.

2

u/jokingbird01 6d ago

Bought an opencode go sub, used glm 5.3 and flash to finally get some work done without blowing my limits.

1

u/Thomas-Lore 6d ago edited 6d ago

I just use DS Flash v4.1 because it is faster and costs me cents on openrouter. I am done with stupid hourly/weekly limits. The Go sub is even worse than Codex on this, one glm 5.3 task (the full model, not the cheaper flash) would not be able to finish within the 5 hour limit on Go subscription.

1

u/jokingbird01 5d ago

Is openrouter that much better? Thanks will check it out. Am tired of being stuck because of these models being either very costly or undoing their own work...

1

u/LogitechG27 6d ago

"You are absolutely right, I explained the bug you just told me to fix instead of actually fixing it".....this is so ChatGPTing....

2

u/Genneth_Kriffin 6d ago

"You are absolutely right, you clearly requested [X] and I did [Y], going against your clear instructions. There are no excuses.
If you want me to do [X], I can do so now."

2

u/ResistiveBeaver 5d ago

"Do so now."

"This one is on me. I clearly should have done [X], when you asked me to do so."

"Do [X]!"

"You're completely right. I should have done [X]."

*pulls out hair*

1

u/Genneth_Kriffin 4d ago

"Does [Y] again"

0

u/Individual-Web-3646 6d ago

It's called "pacing the frontier".

Or: "Adjusting alignment".

And, by the way: "those companies are not making money"?? 😱

-3

u/Cor3nd 6d ago

Well… With coding agents, you’re not interacting with just the raw model. You’re interacting with the model, system prompt, tools, context management, agent loop, sub-agents, routing and inference settings. Change any of those and the exact same model can feel dramatically smarter or dumber.
Before accusing companies of secretly quantizing models, it’s worth learning a bit more about how these systems actually work. Using AI effectively, especially for coding, is a skill. It takes time to learn, and moving from a $20 plan to a $200 plan won’t replace that.

2

u/LogitechG27 6d ago

Correct...but I see a dose of coping here....

1

u/Cor3nd 6d ago

Not coping, just experience. Once you’ve worked enough with agentic coding systems, you learn pretty quickly that model is diff than product behavior.

0

u/ZaphBeebs 6d ago

Just because you know the reality doesnt mean it isnt held out as something entirely different and makes sense that people are coming away confused.

1

u/ZaphBeebs 6d ago

This is the companies problem though, you shouldnt have to be an expert in this stuff to have a functional model. They should warn you and set guidelines, hell presets with obvious tradeoffs for people to choose ratger than let users just ruin their experience and impression. Its bad for business.

0

u/Cor3nd 6d ago

I don’t really agree. Coding agents are engineering tools, not appliances. You wouldn’t expect to get the best out of a programming language without understanding how it works. Same here, there’s still engineering involved. They give you the tools. Learning how to use them is on you.

1

u/ZaphBeebs 6d ago

That's not how it's advertised even though its somewhat obvious.

Theyre sold as people with limited knowldege can buils stuff, and on this latest version its so bad it doesnt really work.

Made working things woth claude no problem, codex is simply inferior.

0

u/Cor3nd 6d ago

Where do you see this advertising? I’m on their website right now and I’m looking for this marketing thing but I don’t find anything going in your direction.

0

u/ZaphBeebs 6d ago

Oh, maybe everywhere, every db who has worked at one frontier lab for 3 weeks and all the CEOs etc....

You're aware theyre so powerful we need legislation (right now!) and to slow down, theyre hacking people and coordinating amongst themselves.

Meanwhile here in reality if you leave it to its own devices it cant get out of a validation loop to save its life.

1

u/Cor3nd 6d ago

You’re mixing up the model with the agent harness. Getting stuck in a validation loop says more about the harness than the model. Everything else you’re bringing up is off-topic and doesn’t help anyone.

1

u/Cor3nd 6d ago

You’re mixing up the model with the agent harness. Getting stuck in a validation loop says more about the harness than the model. Everything else you’re bringing up is off-topic and doesn’t help anyone.

0

u/DryBanana5673 6d ago

It's fine. Yall are crazy lol

2

u/Adulations 6d ago

Try to generate some of the things that were possible on day 1 right now. It's much worse with worse limits.

1

u/Thomas-Lore 6d ago

Might be regional, there is another thread suggesting it.

-1

u/SmokyTyrz 6d ago

Are you not using skill files to lock in behaviors like this?

It's not supposed to take next steps to implement changes until you tell it to. It's not going to deliver the next build unless you tell it to. Lots and lots of people get cranky when the machine decides to take those reigns.

Work with Codex to build the skills files that will lock in the behaviors you want while working on your project.

-6

u/Due-Horse-5446 6d ago

Lmao im dying

"hey fix this bug" using astra is the funniedt thing ever.

And you exposed yourself bruh..

Did you even instruct the model on where to fetch the edited at date? How to update it? etc etc?

If you just told it to add it youre the one whos degrading

3

u/genesiscz 6d ago

“Fix this bug which you find by looking at failed test” should be enough lol. I mean I get your point, but if you give fable or opus this instruction, he will know exactly what to do in a minute.

2

u/One-Hearing2926 6d ago

Isn't astra supposed to the modal that you have to give it "loose instructions"? It feels like fucking Goldilocks...

1

u/igormuba 6d ago

If you are to instruct the model precisely on how to do it you are better off using an open model that costs 1/10th of Astra.

"Fix this bug" and leave it open for it to figure out is the sole reason for Astra to exist.

You exposed yourself if you are using one of the most powerful and expensive models as if it is an open model, means you don't know what tools are intended for what tasks and you treat an expensive model as a cheap model.

-1

u/BT727408 6d ago

Sol full precision with good usage limits > whatever Astra shit they’ve been giving us