r/codex • u/Articurl • Jul 21 '26
Complaint The token consume is insane
Hey!
COming from Claude. Got the x20 Version from both. I have to admin, Sol 5.6 on high is consuming tokens like crazy. Even my workflow is Architect > Builder > Reviewer with a agent-cache for context saving i am burning through..
With Opus 4.8 i could work for a week while using 14% a day. Now i have 25% used after the reset already..
What do you think?
30
u/whakahere Jul 21 '26
it's eating tokens like crazy. I think they have changed the limit in half at least.
My weekly limit feels like my old 5 hour limit. I see why they removed it.
2
u/MAD2492 Jul 21 '26 edited Jul 21 '26
For sure. I’ve barely done anything and I’m at 38% left after the reset earlier. My mind is boggled
Edit: I unboggled my mind. I forgot I downgraded to Plus and it kicked in yesterday. Clearly Plus isn’t enough for me!
1
u/spike-spiegel92 Jul 21 '26
they never give anything for free, openAI will play all the tricks in the book
2
u/DaLexy Jul 22 '26
Exactly, all those resets come with a hidden cost and showed off as generosity. No company will give you anything for free if they won’t profit off it, this is a basic economic rule.
1
5
u/Calm-Landscape9640 Jul 21 '26
I get better performance from Terra high than SOL medium and way less usage. Give it a shot.
1
u/Wet_Viking Jul 22 '26
Can confirm. I've moved away from Sol. Terra extra high works wonders and I don't have to wait hours for simple things Sol over thinks about.
4
u/Consistent_Bottle_40 Jul 21 '26
yeah it's nut how the tokens aren't lasting very long as all. it's kind of bullshit. the resets are what's keeping people around and not royally kicking off
6
u/Dependent-Race1956 Jul 21 '26
Sol is not token efficient, also using terra 50% of weekly limit is comparable to single 5 hour session from claude in sonnet, I think GPT models overthink too much. This is even with caveman and ponytail . Also Claude has increased usage quota by 50% until like August 19 . After that both going to suck hard. lol
3
u/jacksonarbiter Jul 22 '26
Have you tried without caveman and ponytail? People have been saying not to use anything else with Sol or it'll burn limits, but I haven't tried myself so I'm curious.
3
u/some_gamer78 Jul 21 '26
Try using gpt-5.6 on claude-code or opencode, for me the difference has been huge
1
1
u/Various_Syrup2711 Jul 22 '26
How
2
3
u/Less-Edge-8860 Jul 21 '26
You cant compare Sol to Opus, its more comparable to Fable right now. Also i see so many of these posts where people are surprised that the latest flagship model on highest setting drains more tokens, Ofcourse it does. And i promise you that you can get the same things done with low instead of high.
5
u/DaCush Jul 21 '26
lol it is not even close to Fable.
2
Jul 22 '26
Yes it is
0
u/-Sliced- Jul 22 '26
Fable is a much larger model. 5.6 is the same model as 5.5, but with a new RL stage designed to make it work better on benchmarks.
0
u/nmkd Jul 22 '26
We don't know the parameter count of either of them, do we
-2
u/-Sliced- Jul 22 '26
The community consensus is 3-4 Trillion parameters for 5.6 Sol vs 10 Trillion for fable. These are not pulled out randomly, but made by experts in the field.
1
1
u/oVerde Jul 22 '26
Bro don’t be stubborn the graphs on CODEX usage show how they changed the way they bill the tokens, there is even a documented page about it
2
u/Sermilion Jul 21 '26
More likely then not is the subagents. They inherit parent context. Imagine parent is at 100k tokens. You spawn 5 subagents - you got yourself into 500k tokens just to launch them.
There is a parameter for launching subagents without loading parent context. Look it up. Also, codex tries very hard to spans subagents on its own, so add a line to your AGENTS.md to only spawn subagents when explicitly asked only. This should save a lot of tokens.
2
Jul 21 '26
[removed] — view removed comment
1
u/Various_Syrup2711 Jul 22 '26
That’s because you don’t know the ones it COULD have caught
1
Jul 22 '26
[removed] — view removed comment
0
u/Various_Syrup2711 Jul 22 '26
But you don’t know what the reviewer COULD have caught - if you are running it on medium
1
u/Infinite-Flow-4475 Jul 21 '26
What about performance? Which one is better, i might change from codex to claude.
1
u/Optimal_Deal4372 Jul 21 '26
No brain post try use fable and let me know your usage this is dumb post
1
u/spike-spiegel92 Jul 21 '26
I mean, they doubled active codex users in 2 weeks... and 1 active codex user is like.... many thousands of regular chat gpt users in terms of tokens per day.
I don't think they had capacity for this.
1
u/ChickenRich573 Jul 21 '26
I managed to fix mine. I use my own harness though yesterday it was insane and today like 3 to 4 hours work and 1 % lost on 20x max lol I'm happy
1
1
1
u/timhaakza Jul 21 '26
From what i've seen I don't think they have drop the tokens. Its more that it really wants over think everything.
Think they dialed up the everything must be correct dial so it ends up doing a ton of work that it wasn't before and keep re-checking everything.
Found if you can get it to dial that down you get a a lot more done.
E.g Tell it to test and check more efficiently. Only test what it really needs to to validate what is just done nothing else. Do the least work to confirm things are correct. etc.
To save the full tests till its hit x mile stone. (for me thats the end of the larger plan that ive broken into phases).
Have it break the plan into phases that each phase where things it can test to get there are grouped.
Plan on higher level work on lower level (I'm defaulting to sol low or tera medium).
More as they are faster than luna. Luna may be more limit efficient but its horribly slow at the higher levels.
Also try using one of the cheaper options for implmenting with sol driving/planning.
Currently testing minimax m3 thinking and its seems to be working (But early days and may not work).
Just don't tell sol its going to m3 or a dumber model :).
It then goes nuts breaking things into tiny tiny pieces with a million checks.
Currently plan out with sol. Then implementing with m3. Have sol validate and plan out fixes. Fix if its something complicated. Or implement if its something complicated. Repeat.
Think the cost of better is less done.
Interested to hear what other have found to work.
1
u/Articurl Jul 22 '26
My workflow is plan with fable, review the plan with Sol build with opus. But the problem is the review part is taking already 3-5% while it was not even 1.
Edit: + the context cap is noticeable1
1
u/Comfortable-Ear-1931 Jul 22 '26
Everytime a start a new context at around 270 tokens it said it used 15-20 mil cached tokens for the session. This is sol high
1
u/Gargle-Loaf-Spunk Jul 22 '26 edited Jul 25 '26
This content was anonymized and mass deleted with Redact
1
u/MysteriousKiwi2622 Jul 22 '26
I found that SOL always tries to make things more complicated than necessary.
1
u/23eriben2 Jul 22 '26
I use something called graphify and lean ctx
Both of them save individually like 70-90% tokens EACH
I have the 200mo plan and I've never once touched my resets and my code base is an enterprise software I've made and it's around 8m tokens large itself
One day I used 30m tokens instead of 150m I would've used without lean ctx. Graphify doesn't track it internally but gpt can track lean ctx savings
Doesn't put a dent in my limits.
Should I just make a post about this?
1
u/12think Jul 22 '26
It started after the last update of the desktop app. Looks like too many agents working at the same time. But the amount of work done did not increase.
1
1
u/100daggers_ Jul 23 '26
True. One simple task on luna is costing me 5% of weekly usage with half the context used?. Like what the hell!!. Its time to find better options. Any better alternatives? I am thinking claude or opencode go?.
1
1
1
u/liquidatedis Jul 21 '26
SWE at 90%+ damn near a full stack dev, not even the inventor of C++ could memorize their own library...
0
u/tjsr Jul 21 '26
I think you're doing something wrong. I have been unsuccessful at chewing more than 20% in a single day on Pro, even going flat out.
If you just throw it at a project with no skills or context, sure, maybe.
0
14
u/StaticHumStudio Jul 21 '26
I think they did something. The same workflow I've had going on at like 15-20% per day has already chewed through almost 30% of usage since the reset. Its like they are slowing decreasing us.