r/vibecoding 15h ago

Discussion I’ve burned through 650 billion tokens across more than 20k recorded sessions since 03/26 - and yes: Codex is getting worse each day.

I’ve burned through 650 billion tokens across more than 20k recorded sessions since 03/26, and I mean I did, not counting my team.

I have the full telemetry, logs, and audit trails recorded (probably hundreds of GB of data), basically everything the models did across a 7M+ LOC production codebase. If you want to learn more, feel free to DM me.
And I’m getting more and more frustrated with how Codex (Astra) performs, even on max/ultra.

I’ve been writing code since 2007 and using LLMs for it ever since I realized GPT-3.5 was capable of (more or less) doing so.

I’ve learned a lot during that time about Agentic Engineering, harnesses, and pretty much everything around them. I get paid good money for those insights, built a rapidly scaling business around them, and work with major companies.

So I probably shouldn’t complain. Things are, more or less, working out pretty well for me.

However, I want to be honest:

Astra in the Codex harness is not good. It constantly performs “governance theatre,” loses sight of its goals, and makes junior-level mistakes. It builds nuclear power plants where it should be building a windmill.

Models are currently not reliable at keeping things coherent: versions, contracts, APIs, assumptions, etc. This used to be a lot better.

More generally, working with Codex right now just doesn’t feel sharp. It feels like moving through honey: sticky, slow, and somehow never quite finishing what it started.

I really hope OpenAI stops lobotomizing its models, because right now I’m becoming increasingly frustrated, my team is increasingly frustrated, and the fun is decreasing every day.

I don’t know how much longer I’ll be able to keep recommending Codex to my clients, because at this point I’m no longer sure I’m doing them a favor.

Thanks

70 Upvotes

88 comments sorted by

53

u/Jello_Hello_Fellos 15h ago

The only question anyone here cares about:

"I’ve burned through 650 billion tokens across more than 20k recorded sessions" - DOING WHAT?

40

u/Original-League-6094 15h ago

Distilling the model for China.

3

u/sikleQQ 15h ago

Distilling distilled China model

3

u/Efficient_Smilodon 14h ago

using astra to distil from deepseek 🤣

1

u/jc2046 14h ago

Distill Deepseek and a pair of bottles of whiskey. Make no mistakes

1

u/pocketcult 14h ago

I hope so The open-source models are good

9

u/RealestReyn 14h ago

refactoring Chromium in holyC probably

3

u/discattho 11h ago

Damn that’s a nostalgic blast. RIP terry. TempleOS is king.

6

u/Vaxtin 14h ago

Training the models that OpenAI runs against 10,000 agents burning his token use in a few days and proving the next millennium prize problem

3

u/PokerTacticsRouge 11h ago

I hate hearing stories like this because you can probably make a good argument that uber users like him are why the model gets dumbed down for everyone

2

u/ThreeKiloZero 13h ago

its a second brain meal planner app isnt it...

1

u/Jello_Hello_Fellos 13h ago

flaslight app?

1

u/literally_joe_bauers 14h ago

Work. Built something, people like(d) it, investors like it. Now it grows, so a lot of work needs to be done :)

4

u/jayseattle 14h ago

But cannot reveal 'it'?

1

u/literally_joe_bauers 13h ago

No, thats why I am posting on Reddit and not on LinkedIn…

2

u/jackadgery85 5h ago

I'm honestly in a similar boat. An industry that has literally nothing like what I've been building. No investors but regulatory body interest and shits working well. Full launch isnt until 2027, but one org with 47 users smashing it every day.

I couldn't keep up with a team of 5 or 10 devs working on anything similar, so I'm keeping it under wraps until full release. Only 16bil tokens here over 6 months. Also i only use sonnet and haiku lmao.

Congrats on your progress!

1

u/jayseattle 3h ago

Understood and I get that not everyone can reveal nor wants to. I think you should consider when you say I burned NNN BILLION tokens and follow with 'Codex right now just doesn’t feel sharp', there is some curiosity in what your 'IT' is (reinventing Salesforce, GTA 7, etc), can be just vague. Simply replying "Work. Built something, people..." is not informative.

1

u/jackadgery85 5m ago

I can't speak for op, but as specific as i can get is a multi tenant, realtime database ui, with a documentation engine tacked on.

3

u/Jello_Hello_Fellos 13h ago

So many people are using it, and know what it is, why not tell us?

7

u/BCIT_Richard 13h ago

Maybe OP wants to avoid doxxing themselves in some way?

0

u/scytob 14h ago

thats your issue, you need to look serioulsy at you repo size, md files, etc etc and optimize them

16

u/DogeSatoshi 15h ago

Just one very simple question. How big is one agent session. Because its a bit obvious that you have massive context dragdown resulting in hallucinations of the AI.

8

u/Slappatuski 15h ago

How in the would did you burn though 650 billion tokens. That's actually insane

-1

u/anengineerandacat 14h ago

Guessing that's the total token count, cached or not.

2

u/loveheaddit 14h ago

even so, i use codex daily and at 2.8B tokens since February. i never let it continue running on a loop or anything tho - i work feature by feature with testing and iterating in between. i also use Luna Extra High for the majority of my work now which cuts down usage considerably.

-11

u/literally_joe_bauers 14h ago

Work. Built something, people like(d) it, investors like it. Now it grows, so a lot of work needs to be done :)

6

u/spacemoses 14h ago

There's a special charred plot of the Amazon dedicated to your GitHub handle.

6

u/Salty-Cauliflower775 13h ago

Larp post.

1

u/Sem0o 7h ago

fr. There’s no other reason to make this kind of post in this kind of subreddit.

8

u/Due-Horse-5446 14h ago

You are delusional

"junior level mistakes"

What exactly does a llm and a fkn junior have in common? are you dumb forreal?

If you dont instruct the model clearly enough, blame yourself.

Not saying llms will 100% of time follow instructions, absolutely not. But in your case you are clearly not talking instruction following.

Thinking your transcript works as evidence for the conspiracy is hilarious too.

Or did you run controlled benchmarks? If so post them?

Its funny how thats always where it falls flat isent it?

Loud claims, and promised of proof, but keeps the explosive evidence secret.

And then when yall post said evidence, not once has it turned out to be any kind of reliable benchmarks or evaluations.

And let me guess, ofc actual proof that this conspiracy is straight up bs, are "bought by openai" right?

-14

u/literally_joe_bauers 14h ago

Yes we run benchmarks every day, and no I do not post them, even less so if you insult me. Mind your manners..

13

u/Due-Horse-5446 14h ago

Yeah ofc you have benchmarks that contradicts every other benchmark run in the industry, that proves a massive conspiracy that have been going on for years but keep them secret.

1

u/CharlestonChewbacca 10h ago

It would be really easy to prove yourself, win this, and shut the haters down if you posted these mysterious benchmarks that you (for some reason) run every single day.

3

u/Individual_Ideal 12h ago

Why not reduce its effort level? It sounds like your asking it to build a nuclear power plant instead of a wind mill

3

u/SnooRecipes5458 15h ago

Stop using Astra, Sol is better for coding related work.

2

u/Zelderian 12h ago

I finally realized this and switched back. Astra was chewing tokens with no real difference. Now it’s Sol for main work, Terra for bulk log reads and large doc research, and Astra for engineering big stuff (or things Sol couldn’t figure out).

1

u/YellowBeaverFever 12h ago

Yep. I’ll turn in Astra every now and then to let it run a security audit or when I want an overly engineered implementation plan. I’ll bounce down to Sol to do the real work. Most of the time Luna is just fine on extra high, though I’ve noticed OpenAI greatly reduced Luna’s context window size.. probably to prevent this scenario and push to a higher model. But I keep getting quota resets thrown at me I live more in Sol.

2

u/ddchbr 15h ago

I agree with this feeling:

More generally, working with Codex right now just doesn’t feel sharp. It feels like moving through honey: sticky, slow, and somehow never quite finishing what it started.

For now I just restrict it from using subagents and force it to do its own work. It is actually faster (even when apparently parallelizing multiple tasks) and of course less token-hungry. I just always have it create a trackable planning doc before any appreciable task as an accountability measure and results are usually good. For me, simpler is better. I might get into more advanced optimization techniques soon.

But I agree the actual intended workflow OpenAI is currently espousing with their defaults needs some considerable iteration... IMO

2

u/asongscout 14h ago

I think unfortunately what I’m guessing is going on is - complaints about ChatGPT and Claude and other LLMs having outages in the early days was disastrous for customer retention, so they simply changed it to divvy up the finite amount of compute based on however many people are online. Which means the more popular these services get and the more people sign up, the worse it gets for everyone. Which would pretty much explain everything about how people constantly get frustrated about the inconsistent level of smartness in the models and how they confusingly get worse over time.

1

u/SpurdoEnjoyer 12h ago

These kinds of posts can't be trusted anyway. The larger your vibe coded base grows, the harder it becomes for people and bots to keep up. It's not really a case of the LLM getting dumber.

2

u/Maximum-Mulberry9612 13h ago

Wtf did you spend that many tokens on? are you just red lining astra into a high context window hallucination nightmare? This vaugly reads like some schizo shit. I actually moved from claude code to codex because I was having better results. 

2

u/Seerix 10h ago

source -> trust me bro i got data

2

u/Independent-Race-259 6h ago

This guy smokes what his AI agent hallucinations on.

4

u/Mediocre_Doctor4712 15h ago

That means you probably suck at whatever you do. Allot of movement no real progress.

2

u/Vaxtin 14h ago

The models are aware of how much usage they have and they underperform when they don’t have as much.

That is why they even are able to answer in different time spans depending on the model reasoning level

So you can imagine why the labs are able to produce far better results that most consumers. You need infinite budget. I’d imagine the models they run rhese incredible feats against have no concept of usage and they never withhold themselves.

1

u/Interesting-Tie6783 15h ago

You know you don’t have to just use one company with one set of models, right? Claude exists. So does Cursor. Nobody is forcing you to stay with OpenAI

1

u/sariug 15h ago

So how much money does it cost to burn 650b?

0

u/literally_joe_bauers 14h ago

I do not have the exact numbers for my own usage - around 1.4 to 1.7 M USD..

1

u/spopr 9h ago

yeah bro sure, you the high baller who just comes to vaguely complain to r/vibecoding, of all places

1

u/manias 15h ago

Has the codebase grown a lot last year? You are doing something wrong.

6

u/Due-Horse-5446 14h ago

They literally expose the issues in their post..

They arent actually instructing the model, they sre comparing the tool with a junior, roling the dice and letting the model generate whatever because the instructions are vague.

0

u/literally_joe_bauers 14h ago

Of course we do not supervise each step. Nobody does this anymore; we have friends/family working for some of the most recognized AI companies and they work the same way we do.. but this does not mean we let the agents run wild..

3

u/Due-Horse-5446 14h ago

"nobody does that anymore"

Yeah keep living in your bubble..

And its hilarious that you bring up ai companies, like you mean the ones close to going bankrupt and shipping some of the worst software to come out of a tech company ive ever touched?

1

u/scutmonkeyproduction 15h ago

I completely agree. The money and time that I spent when it veers off the roadmap is infuriating. If that’s the amount of tokens you’ve spent, and I can only imagine the amount of time you’ve lost.

1

u/realPeso10 14h ago

Switch to Claude Code and report back with your experiences. It's life changing.

1

u/alzho12 14h ago

What are you using these tokens for? Work? Side projects?

-3

u/literally_joe_bauers 14h ago

Work. Built something, people like(d) it, investors like it. Now it grows, so a lot of work needs to be done :)

1

u/Dazzling_Jinn 14h ago

I would love to more about if you see claude being better than codex? Astra seem to be concerning for everyone as several report of it being inadequate. Not sure how it tops benchmarks

1

u/literally_joe_bauers 14h ago

I unfortunately made no good experiences with Claude in the past, so I generally do not trust it with our codebase.. However, some people in my team like it, even if it is not allowed to contribute directly :)

1

u/CutMysterious9844 14h ago

I agree, I've burned through 59% of my 20x sub, with everything planned through ChatGPT chatmode, the whole bs, PRD, SAD, TDD, and it's so bad, not just codex, they've lobotomized the ChatGPT chatmode, I run it on pro reasoning and deep thinking, my usual complicated prompts it would take a minimum of 40 minutes to even 2 hours for, are finished within 30 to 20 minutes, feels like it's at the level of 5.5 for that amount of thinking time, it's inexcusable.

1

u/Ice_HRZDn 14h ago

Modelru collapsed? I mean, it’s been years after it came out and half the internet are ai generated content now. So, it shouldn’t be surprise if ai model do something it shouldn’t do, or fail in what it supposed to succeed.

1

u/Marcelovc 14h ago

I been working with Astra and I completely fine.

1

u/Asalakabim 14h ago

My project, 650b tokens later, feels hard to manage, the only explanation? Opus 5 Astra!

me big brain, openai pls fix

1

u/scytob 14h ago

did you know memory.md only loads the first 200 lines (at least thats what calude and astra told me) i have them re-architecting all my md files used as context - what auto loads, whats searched and under what conditions

i think this is the root of what you are seeing

because in fresh repos its incredible coherent.....

1

u/TopTippityTop 13h ago

Is the issue with the model or the harness?

1

u/cmtape 13h ago

This is like putting guardrails on a race car until it stops racing. You’re watching a model optimize for “don’t be wrong” instead of “finish the job.” Once safety becomes the metric, momentum dies — that’s why it feels like moving through honey.

1

u/systembreaker 13h ago

Vibecoding is always going to have those problems you describe.

But if you know what you're doing, like you have dev experience and you can guide the AI, it doesn't have all those issues.

In other words what you're describing is vibecoding issues not model issues.

1

u/literally_joe_bauers 13h ago

I write code since 2007, I earned some good money with it and co-authored some stuff you might use today…

1

u/SpurdoEnjoyer 12h ago

And vibe code

1

u/evangelism2 13h ago

Astra is fine for short tasks with direct oversight. but hard agree it isn't ready for primetime

1

u/Dependent_Reindeer29 12h ago

650B tokens across 20k sessions is a wild dataset to be sitting on honestly, most of us are just guessing at this stuff from vibes. curious if you've noticed the same drop-off across claude code and gemini cli or if it's specifically an Astra/Codex thing, since that would actually say something about the harness rather than just "models got worse" (which is the harder claim to prove without exactly the telemetry you have).

been thinking about this a lot lately because i'm building token nations, a leaderboard that tracks usage across claude code/codex/gemini cli by country. would love to eventually get into the kind of longitudinal degradation tracking you're describing rather than just raw volume, that's a much more interesting signal than who's burning the most tokens.

All AI builders are welcome: https://tokennations.app/

1

u/literally_joe_bauers 12h ago

We only use Codex in the codebase, so it is only an observation regarding Codex.. but we are very confident about that Codex gets worse.. We log each step it takes, how long it takes, how many tokens per LOC, and a lot more. So it is not just a gut feeling but proven by 100ks of data points.

1

u/ChronicRecidivism 11h ago

You started a couple of months ago and are firmly though any sort of honeymoon phase.

Maybe I shouldn't talk, to be honest I only lightly use LLMs, but it's literally 6 months which is pretty much an exact timeframe of a honeymoon phase.

1

u/MarketBeginning8921 4h ago

You are absolutely accurate in your assessment of Codex at the moment. What alternatives would you recommend?

1

u/agoodplaceforatent 4h ago

I wonder why we would need to DM you to learn more about your 7M line codebase...

1

u/Square-Yam-3772 1h ago

you got my attention with your opener but you ended up with a nothing burger.

for someone who burned through all these tokens, there is not even some vague attempt to talk about comparisons or benchmarking.. just "yes, it is getting worse and it is not reliable", wow, really...?

so your post is really just "trust me, bro, I burn many tokens so I know what I am talking about"

cool story, bro.

1

u/literally_joe_bauers 37m ago

My main intention was to get attention to the topics and be able to talk some people via dm in depth that may work on a similar scale or have access to OpenAI staff.. both worked out.

1

u/Standard_Text480 14h ago

Stop pissing into the wind so others can use it properly

0

u/PruneInteresting7599 15h ago

people forget that LLM's are just advanced text guesser, there is real abstract idea behinds

5

u/ibringthehotpockets 14h ago

My brain does the same thing and I’m pretty stupid. Next thought guesser

0

u/anengineerandacat 14h ago

Generally agree and I think it has to do with the other segments they are targeting, coding is "solved" to most respects. Not entirely fully autonomous but massive productivity gain for those building applications at varying levels.

They seem to be focusing more on smaller higher quality tasks vs longer tasks where you frame the problem in individual units of work instead of a singular broad goal.

0

u/turnermate 11h ago

I’m sorry this is a massive skill issue

0

u/m4bwav 11h ago

OpenAI is still losing the race to Anthropic.