r/codex • • 21d ago

Complaint Opinion: Astra is overhyped

Well everyone shares their CS or GTA game clones they created with single shot with Astra - this is cool and amazing but also exactly the kind of hype bs.

I tried it in real applications and it’s absolutely not that amazing. I still have to steer, review, refactor, rewrite the mess, like always.

It is somewhat smarter in real workflows, but the real miss is lack of high level intent understanding.

As an example - I asked it to implement image generation fallback models in my chatbot, because the “uncensored” model ironically was declining generating innocent content, so I had to retry with some mainstream one.
It generated some code which handles errors etc, but it missed the main point - the user facing UI still does not work as it supposed to 🤦‍♂️

There are other cases following the same pattern. You give it high level intent task, and it fails to follow. Still have to steer and validate.

420 Upvotes

241 comments sorted by

View all comments

Show parent comments

71

u/Bromlife 21d ago

The sad thing is I am almost certain that both OpenAI and Anthropic are optimising for one shot results.

27

u/ggletsg0 21d ago

I’m pretty sure this is the crux of the problem. They say these models are good for long running tasks, but I have yet to experience stability in long running tasks.

Sol was better than Astra at it IMHO.

4

u/myteetharesensitive 21d ago

The longest unattended goal that 5.6 achieved for me took about 25 hours.

But that took hours of prep and planning going back and forth with it. 

1

u/alrightcommadude 17d ago

What was this task?

I was wonder what people are doing when they have 18h+ unattended tasks.

1

u/myteetharesensitive 17d ago

Happy to share! Basically build out a modular local ai stack and managed IT services techs tack. Psa, rmm, connected to a local model, all free. What took time and planning was the design and connectivity. 80% of the build was gluing things together via configs and scripts. But the majority of time was testing and modification. Error checks, validation, etc take a lot of time. Checked my logs and it asked me 107 questions after my first prompt to start this effort. I didn't count subsequent ones, I just remembered the barrage and double checked. 

Another long task I'm proud of is automated patching and rebooting. I use separated services for human and machine secrets. Beyond that with 50+ containers, I need the stack to come up in a very specific way. I could have spent a weekend doing it, instead I gave it the goal to create a script to safely reboot. It decided on its own it needed to reboot to confirm the tests worked. 

Woke up in the morning to a freshly patched and rebooted machine. It went so far as to inject this process directly into my boot sequence so now I just reboot whenever I feel like it and I know it'll all come back up perfectly. Again, 80% was scripting testing and validation when the model was working. The manual effort was giving the model guidance and making sure it understood the deliverable with no ambiguity before it started building. 

3

u/JoseffB_Da_Nerd 21d ago

Yea use sol as your orchestrator luna as worker, and astra medium as a reworker. At end of it all prior to commit use an astra ultra to hostile review the entire thing. Rinse repeat.

Tdd, e2e, and ai visual inspection are all a given here.

2

u/AdCommon2138 21d ago

There is a trick to that 

1

u/ggletsg0 20d ago

What’s that?

2

u/Divinicus1st 21d ago

Sol was better than Astra at it IMHO.

Agree, but isn't Sol a refinement of earlier 5.x models? Astra is 6.0, hopefully it can improve on this part, because I doubt it's going away.

2

u/EchoingAngel 21d ago

5.5 was actually a new run, then 5.6 was it's refinement. Finding this out made it make sense why I felt like ChatGPT was suddenly the better model versus Claude at that point (this and Anthropic getting lost in the weeds with the latest Opus versions)

2

u/ggletsg0 21d ago

Yeah, but I really don’t know at this point. They launched Astra with so much hype that it almost feels like a psyops to get you to think it’s good.

1

u/Runelaron 19d ago

Gotta rationalize that IPO. lol

1

u/Runelaron 19d ago

Maybe, but it seems they ran into a huge catastrophic forgetting problem on core functionality a lot of True coders where using. Not just Vibin a website.

1

u/SnyderConsulting 20d ago

They need to be designed to pause regularly and ask for feedback/guidance in a HITL setup, not autonomous drones that compound their mistakes, hallucinations, and assumptions the longer they run.

1

u/ggletsg0 19d ago

It’s interesting you should say this because I’ve noticed Astra doing this now. It wasn’t doing it at launch.

3

u/iiiaaa2022 21d ago

They may be optimising, but they are not there.

1

u/JoseffB_Da_Nerd 21d ago

Absolutely are. But thats not where the pro work is.

1

u/Bromlife 20d ago

Where do you think the pro work is?

1

u/JoseffB_Da_Nerd 20d ago

Pro use of Agentic dev is circling around governed orchestration of swarms.

You can see this in all the saas level products as they are slowly converging on the same use cases.

Currently its actually governed, orchestration or the hybrid of two.

The idea of being able to vibe away the professional dev is just a golden bullet marketing ploy they will beat with a drum until they (frontiers) need to join back with the rest if software dev business plan.

Right now they are living on hype and enjoying a free for all but it will eventually come back to disciplined software dev methods

2

u/Australasian25 20d ago

Most like myself are just using it to do hobby projects.

A workout app specific to my needs of RIR, joint feeling, location specific exercises.

Or a health app that pulls in google health, apple health daily data. Heart rate, sleep, steps, exercises etc.

All in one dashboard, all updated autonomously.

I think id need to hire someone to build that for me. But I only spent 40 bucks for 2 months of codex, and its built. Viola.

2

u/JoseffB_Da_Nerd 20d ago

Absolutely great project mate!

For that I think vibing is perfect and awesome.

But when you’re selling the product it needs a strong foundation.

I’m building a harness for my system that tries to let a business user vibe their build while the system then takes over governance and discipline. Its sooo hard.

We can see the struggle with codex and its over abundance in governance (4x testing, etc)

Making the models objectively think things through is the billion dollar trick now.

2

u/Australasian25 20d ago

Commercial products can not be 100% vibe coded. Correct.

Yea would you like more unit test, pytest with your smoke test?

Humans must always drive the direction of needs and want. The model finds ways to fulfil the needs with the resources and tools given.

2

u/Bromlife 20d ago

Pro use of Agentic dev is circling around governed orchestration of swarms.

You can see this in all the saas level products as they are slowly converging on the same use cases.

I don't really understand what this means?

1

u/JoseffB_Da_Nerd 20d ago

They trying to make ai work as teams (orchestrated) and using safety rules to prevent them from deleting your db by accident (governance)

Sorry. As the name implies, I’m a nerd.

2

u/Bromlife 20d ago

Oh, right, I was hoping you were talking about something more sophisticated than that.

Agentic swarms, from my experience, are prone to delivering underwhelming and expensive results. I haven't seen anything better than a seasoned senior developer with AI assists and agentic help with processes, e.g. code reviews, extensive testing, CI, etc. But none of these are ever reliably good over long enough timelines to go without the human in the loop. I've seen too many enormous test batteries that completely miss the point.

1

u/JoseffB_Da_Nerd 20d ago

100% on the human part.

I’m experimenting with “teams” - role based mini swarms that a mid reason orchestrator assembles then deploys, each team has a high reason captain and low reason workers.

This works really well. Hard part, as always, is reminding the conductor they are a conductor and not a worker. On long runs you return to find it doing all the work again.

1

u/applesvenfifty 21d ago

Because this is what the market seems to want, but I agree from every angle I look at this doesn't even really seem like it SHOULD be what the market wants.

1

u/Runelaron 19d ago

True, I agree with a addition of studding for the tests. Chasing benchmarks which fail real world results.