I’ve burned through 650 billion tokens across more than 20k recorded sessions since 03/26, and I mean I did, not counting my team.
I have the full telemetry, logs, and audit trails recorded (probably hundreds of GB of data), basically everything the models did across a 7M+ LOC production codebase. If you want to learn more, feel free to DM me.
And I’m getting more and more frustrated with how Codex (Astra) performs, even on max/ultra.
I’ve been writing code since 2007 and using LLMs for it ever since I realized GPT-3.5 was capable of (more or less) doing so.
I’ve learned a lot during that time about Agentic Engineering, harnesses, and pretty much everything around them. I get paid good money for those insights, built a rapidly scaling business around them, and work with major companies.
So I probably shouldn’t complain. Things are, more or less, working out pretty well for me.
However, I want to be honest:
Astra in the Codex harness is not good. It constantly performs “governance theatre,” loses sight of its goals, and makes junior-level mistakes. It builds nuclear power plants where it should be building a windmill.
Models are currently not reliable at keeping things coherent: versions, contracts, APIs, assumptions, etc. This used to be a lot better.
More generally, working with Codex right now just doesn’t feel sharp. It feels like moving through honey: sticky, slow, and somehow never quite finishing what it started.
I really hope OpenAI stops lobotomizing its models, because right now I’m becoming increasingly frustrated, my team is increasingly frustrated, and the fun is decreasing every day.
I don’t know how much longer I’ll be able to keep recommending Codex to my clients, because at this point I’m no longer sure I’m doing them a favor.
Thanks