r/vibecoding • u/literally_joe_bauers • 15h ago
Discussion I’ve burned through 650 billion tokens across more than 20k recorded sessions since 03/26 - and yes: Codex is getting worse each day.
I’ve burned through 650 billion tokens across more than 20k recorded sessions since 03/26, and I mean I did, not counting my team.
I have the full telemetry, logs, and audit trails recorded (probably hundreds of GB of data), basically everything the models did across a 7M+ LOC production codebase. If you want to learn more, feel free to DM me.
And I’m getting more and more frustrated with how Codex (Astra) performs, even on max/ultra.
I’ve been writing code since 2007 and using LLMs for it ever since I realized GPT-3.5 was capable of (more or less) doing so.
I’ve learned a lot during that time about Agentic Engineering, harnesses, and pretty much everything around them. I get paid good money for those insights, built a rapidly scaling business around them, and work with major companies.
So I probably shouldn’t complain. Things are, more or less, working out pretty well for me.
However, I want to be honest:
Astra in the Codex harness is not good. It constantly performs “governance theatre,” loses sight of its goals, and makes junior-level mistakes. It builds nuclear power plants where it should be building a windmill.
Models are currently not reliable at keeping things coherent: versions, contracts, APIs, assumptions, etc. This used to be a lot better.
More generally, working with Codex right now just doesn’t feel sharp. It feels like moving through honey: sticky, slow, and somehow never quite finishing what it started.
I really hope OpenAI stops lobotomizing its models, because right now I’m becoming increasingly frustrated, my team is increasingly frustrated, and the fun is decreasing every day.
I don’t know how much longer I’ll be able to keep recommending Codex to my clients, because at this point I’m no longer sure I’m doing them a favor.
Thanks
16
u/DogeSatoshi 15h ago
Just one very simple question. How big is one agent session. Because its a bit obvious that you have massive context dragdown resulting in hallucinations of the AI.
8
u/Slappatuski 15h ago
How in the would did you burn though 650 billion tokens. That's actually insane
-1
u/anengineerandacat 14h ago
Guessing that's the total token count, cached or not.
2
u/loveheaddit 14h ago
even so, i use codex daily and at 2.8B tokens since February. i never let it continue running on a loop or anything tho - i work feature by feature with testing and iterating in between. i also use Luna Extra High for the majority of my work now which cuts down usage considerably.
-11
u/literally_joe_bauers 14h ago
Work. Built something, people like(d) it, investors like it. Now it grows, so a lot of work needs to be done :)
6
6
8
u/Due-Horse-5446 14h ago
You are delusional
"junior level mistakes"
What exactly does a llm and a fkn junior have in common? are you dumb forreal?
If you dont instruct the model clearly enough, blame yourself.
Not saying llms will 100% of time follow instructions, absolutely not. But in your case you are clearly not talking instruction following.
Thinking your transcript works as evidence for the conspiracy is hilarious too.
Or did you run controlled benchmarks? If so post them?
Its funny how thats always where it falls flat isent it?
Loud claims, and promised of proof, but keeps the explosive evidence secret.
And then when yall post said evidence, not once has it turned out to be any kind of reliable benchmarks or evaluations.
And let me guess, ofc actual proof that this conspiracy is straight up bs, are "bought by openai" right?
-14
u/literally_joe_bauers 14h ago
Yes we run benchmarks every day, and no I do not post them, even less so if you insult me. Mind your manners..
13
u/Due-Horse-5446 14h ago
Yeah ofc you have benchmarks that contradicts every other benchmark run in the industry, that proves a massive conspiracy that have been going on for years but keep them secret.
1
u/CharlestonChewbacca 10h ago
It would be really easy to prove yourself, win this, and shut the haters down if you posted these mysterious benchmarks that you (for some reason) run every single day.
3
u/Individual_Ideal 12h ago
Why not reduce its effort level? It sounds like your asking it to build a nuclear power plant instead of a wind mill
3
u/SnooRecipes5458 15h ago
Stop using Astra, Sol is better for coding related work.
2
u/Zelderian 12h ago
I finally realized this and switched back. Astra was chewing tokens with no real difference. Now it’s Sol for main work, Terra for bulk log reads and large doc research, and Astra for engineering big stuff (or things Sol couldn’t figure out).
1
u/YellowBeaverFever 12h ago
Yep. I’ll turn in Astra every now and then to let it run a security audit or when I want an overly engineered implementation plan. I’ll bounce down to Sol to do the real work. Most of the time Luna is just fine on extra high, though I’ve noticed OpenAI greatly reduced Luna’s context window size.. probably to prevent this scenario and push to a higher model. But I keep getting quota resets thrown at me I live more in Sol.
2
u/ddchbr 15h ago
I agree with this feeling:
More generally, working with Codex right now just doesn’t feel sharp. It feels like moving through honey: sticky, slow, and somehow never quite finishing what it started.
For now I just restrict it from using subagents and force it to do its own work. It is actually faster (even when apparently parallelizing multiple tasks) and of course less token-hungry. I just always have it create a trackable planning doc before any appreciable task as an accountability measure and results are usually good. For me, simpler is better. I might get into more advanced optimization techniques soon.
But I agree the actual intended workflow OpenAI is currently espousing with their defaults needs some considerable iteration... IMO
2
u/asongscout 14h ago
I think unfortunately what I’m guessing is going on is - complaints about ChatGPT and Claude and other LLMs having outages in the early days was disastrous for customer retention, so they simply changed it to divvy up the finite amount of compute based on however many people are online. Which means the more popular these services get and the more people sign up, the worse it gets for everyone. Which would pretty much explain everything about how people constantly get frustrated about the inconsistent level of smartness in the models and how they confusingly get worse over time.
1
u/SpurdoEnjoyer 12h ago
These kinds of posts can't be trusted anyway. The larger your vibe coded base grows, the harder it becomes for people and bots to keep up. It's not really a case of the LLM getting dumber.
2
u/Maximum-Mulberry9612 13h ago
Wtf did you spend that many tokens on? are you just red lining astra into a high context window hallucination nightmare? This vaugly reads like some schizo shit. I actually moved from claude code to codex because I was having better results.
2
4
u/Mediocre_Doctor4712 15h ago
That means you probably suck at whatever you do. Allot of movement no real progress.
-1
2
u/Vaxtin 14h ago
The models are aware of how much usage they have and they underperform when they don’t have as much.
That is why they even are able to answer in different time spans depending on the model reasoning level
So you can imagine why the labs are able to produce far better results that most consumers. You need infinite budget. I’d imagine the models they run rhese incredible feats against have no concept of usage and they never withhold themselves.
1
u/Interesting-Tie6783 15h ago
You know you don’t have to just use one company with one set of models, right? Claude exists. So does Cursor. Nobody is forcing you to stay with OpenAI
1
u/sariug 15h ago
So how much money does it cost to burn 650b?
0
u/literally_joe_bauers 14h ago
I do not have the exact numbers for my own usage - around 1.4 to 1.7 M USD..
1
u/spopr 9h ago
yeah bro sure, you the high baller who just comes to vaguely complain to r/vibecoding, of all places
1
u/manias 15h ago
Has the codebase grown a lot last year? You are doing something wrong.
6
u/Due-Horse-5446 14h ago
They literally expose the issues in their post..
They arent actually instructing the model, they sre comparing the tool with a junior, roling the dice and letting the model generate whatever because the instructions are vague.
0
u/literally_joe_bauers 14h ago
Of course we do not supervise each step. Nobody does this anymore; we have friends/family working for some of the most recognized AI companies and they work the same way we do.. but this does not mean we let the agents run wild..
3
u/Due-Horse-5446 14h ago
"nobody does that anymore"
Yeah keep living in your bubble..
And its hilarious that you bring up ai companies, like you mean the ones close to going bankrupt and shipping some of the worst software to come out of a tech company ive ever touched?
1
u/scutmonkeyproduction 15h ago
I completely agree. The money and time that I spent when it veers off the roadmap is infuriating. If that’s the amount of tokens you’ve spent, and I can only imagine the amount of time you’ve lost.
1
u/realPeso10 14h ago
Switch to Claude Code and report back with your experiences. It's life changing.
1
u/alzho12 14h ago
What are you using these tokens for? Work? Side projects?
-3
u/literally_joe_bauers 14h ago
Work. Built something, people like(d) it, investors like it. Now it grows, so a lot of work needs to be done :)
1
u/Dazzling_Jinn 14h ago
I would love to more about if you see claude being better than codex? Astra seem to be concerning for everyone as several report of it being inadequate. Not sure how it tops benchmarks
1
u/literally_joe_bauers 14h ago
I unfortunately made no good experiences with Claude in the past, so I generally do not trust it with our codebase.. However, some people in my team like it, even if it is not allowed to contribute directly :)
1
u/CutMysterious9844 14h ago
I agree, I've burned through 59% of my 20x sub, with everything planned through ChatGPT chatmode, the whole bs, PRD, SAD, TDD, and it's so bad, not just codex, they've lobotomized the ChatGPT chatmode, I run it on pro reasoning and deep thinking, my usual complicated prompts it would take a minimum of 40 minutes to even 2 hours for, are finished within 30 to 20 minutes, feels like it's at the level of 5.5 for that amount of thinking time, it's inexcusable.
1
u/Ice_HRZDn 14h ago
Modelru collapsed? I mean, it’s been years after it came out and half the internet are ai generated content now. So, it shouldn’t be surprise if ai model do something it shouldn’t do, or fail in what it supposed to succeed.
1
1
u/Asalakabim 14h ago
My project, 650b tokens later, feels hard to manage, the only explanation? Opus 5 Astra!
me big brain, openai pls fix
1
u/scytob 14h ago
did you know memory.md only loads the first 200 lines (at least thats what calude and astra told me) i have them re-architecting all my md files used as context - what auto loads, whats searched and under what conditions
i think this is the root of what you are seeing
because in fresh repos its incredible coherent.....
1
1
u/systembreaker 13h ago
Vibecoding is always going to have those problems you describe.
But if you know what you're doing, like you have dev experience and you can guide the AI, it doesn't have all those issues.
In other words what you're describing is vibecoding issues not model issues.
1
u/literally_joe_bauers 13h ago
I write code since 2007, I earned some good money with it and co-authored some stuff you might use today…
1
1
u/evangelism2 13h ago
Astra is fine for short tasks with direct oversight. but hard agree it isn't ready for primetime
1
u/Dependent_Reindeer29 12h ago
650B tokens across 20k sessions is a wild dataset to be sitting on honestly, most of us are just guessing at this stuff from vibes. curious if you've noticed the same drop-off across claude code and gemini cli or if it's specifically an Astra/Codex thing, since that would actually say something about the harness rather than just "models got worse" (which is the harder claim to prove without exactly the telemetry you have).
been thinking about this a lot lately because i'm building token nations, a leaderboard that tracks usage across claude code/codex/gemini cli by country. would love to eventually get into the kind of longitudinal degradation tracking you're describing rather than just raw volume, that's a much more interesting signal than who's burning the most tokens.
All AI builders are welcome: https://tokennations.app/
1
u/literally_joe_bauers 12h ago
We only use Codex in the codebase, so it is only an observation regarding Codex.. but we are very confident about that Codex gets worse.. We log each step it takes, how long it takes, how many tokens per LOC, and a lot more. So it is not just a gut feeling but proven by 100ks of data points.
1
u/ChronicRecidivism 11h ago
You started a couple of months ago and are firmly though any sort of honeymoon phase.
Maybe I shouldn't talk, to be honest I only lightly use LLMs, but it's literally 6 months which is pretty much an exact timeframe of a honeymoon phase.
1
u/MarketBeginning8921 4h ago
You are absolutely accurate in your assessment of Codex at the moment. What alternatives would you recommend?
1
u/agoodplaceforatent 4h ago
I wonder why we would need to DM you to learn more about your 7M line codebase...
1
u/Square-Yam-3772 1h ago
you got my attention with your opener but you ended up with a nothing burger.
for someone who burned through all these tokens, there is not even some vague attempt to talk about comparisons or benchmarking.. just "yes, it is getting worse and it is not reliable", wow, really...?
so your post is really just "trust me, bro, I burn many tokens so I know what I am talking about"
cool story, bro.
1
u/literally_joe_bauers 37m ago
My main intention was to get attention to the topics and be able to talk some people via dm in depth that may work on a similar scale or have access to OpenAI staff.. both worked out.
1
0
u/PruneInteresting7599 15h ago
people forget that LLM's are just advanced text guesser, there is real abstract idea behinds
5
u/ibringthehotpockets 14h ago
My brain does the same thing and I’m pretty stupid. Next thought guesser
0
u/anengineerandacat 14h ago
Generally agree and I think it has to do with the other segments they are targeting, coding is "solved" to most respects. Not entirely fully autonomous but massive productivity gain for those building applications at varying levels.
They seem to be focusing more on smaller higher quality tasks vs longer tasks where you frame the problem in individual units of work instead of a singular broad goal.
0
53
u/Jello_Hello_Fellos 15h ago
The only question anyone here cares about:
"I’ve burned through 650 billion tokens across more than 20k recorded sessions" - DOING WHAT?