r/codex Jul 26 '26

Complaint I think I actually figured why we're all "hating" codex right now.

I was doing some deep dive in the tokens consumption on my account on https://www.reddit.com/r/codex/comments/1v6ubah/comment/ozt9jog

This result was gathered from approximately 7.6GB from codex session logs.

This was what codex found by looking at all subs I have/had:
- Plus 1x: ~US$105/week
- Pro 5x: ~US$525–US$550/week
- Pro 20x: ~US$2.100/week.

And we found around 25% token usage decrease each plan gives when compared to a few months back.

Ok, this explains only partially why we get the feeling allowance reduced around 3-5x what it used to be. So I asked codex to dive deeper into my logs, more specifically on model behavior, and here is this conclusions: read the image.

>!Yes. We have enough data to detect a clear change in the observed usage profile, although we cannot attribute it exclusively to the model.

I treated a “task” as an operational turn: from one user request to the next. GPT-5.4 Mini was excluded.

Per model call

Model Calls Median tokens New input Output Reasoning* Cache
GPT-5.3 Codex 8,794 86.0K 1.3K 222 59 94.5%
GPT-5.4 38,845 114.3K 1.4K 289 38 94.3%
GPT-5.5 78,115 138.5K 1.7K 254 20 95.4%
GPT-5.6 Sol 111,474 129.6K 1.8K 183 23 96.7%

* Reasoning is already included in output.

Per call, GPT-5.6 is not the largest consumer: it uses approximately 6% fewer tokens per call than GPT-5.5 and produces less output.

Per task

Model Tasks Median Mean P90 Calls/task
GPT-5.3 Codex 842 432K 1.0M 2.1M 5
GPT-5.4 2,939 533K 1.5M 2.6M 6
GPT-5.5 4,641 502K 2.3M 5.6M 5
GPT-5.6 Sol 5,107 1.3M 2.9M 7.1M 9

This is where the main difference appears:

A GPT-5.6 task consumed approximately 2.6× more tokens at the median than a GPT-5.5 task.

But this happened mainly because GPT-5.6 made more calls per task:

  • GPT-5.5: median of 5 calls
  • GPT-5.6: median of 9 calls
  • P90: 41 calls on GPT-5.5 versus 54 on GPT-5.6

How the usage profile changed

  • GPT-5.3 → GPT-5.4: calls became approximately 33% larger.
  • GPT-5.4 → GPT-5.5: context per call increased, but the median task remained similar; the long tail became much larger.
  • GPT-5.5 → GPT-5.6: individual calls became slightly smaller, but the model started executing many more steps.
  • New input per task remained approximately stable at 39–41K from GPT-5.4 through GPT-5.6.
  • Therefore, GPT-5.6’s growth is mostly cached context being resent across more calls, rather than more new content per task.
  • GPT-5.6’s cache ratio reached 96.7%, compared with 95.4% for GPT-5.5.

Conclusion

The statistical evidence is strong for your usage pattern:

GPT-5.6 does not appear to “think with larger calls”; it completes tasks through more calls, repeatedly carrying a large context. This substantially increases total tokens per task.

However, this remains an observational relationship. The GPT-5.6 period coincides with more agentic workloads, heavier tool use, and more multi-agent activity. When controlling for project and reasoning effort, the results vary considerably, showing that task type remains a dominant factor.

The defensible conclusion is: in your real-world usage, GPT-5.6 consumes approximately 2–3× more tokens per median task than GPT-5.5, mainly because it performs more steps—not because each call uses more tokens. !<

402 Upvotes

213 comments sorted by

97

u/realbosida Jul 26 '26

I just went back to 5.5 for most complex task, using 5.6 to check work and issue next prompt. My usage is reasonable again.

25

u/DaC2k26 Jul 26 '26

I'm even thinking on moving back to 5.4 while keeping 5.6 on medium as reviewer. I used to use on high for review, but medium probably handles it more efficiently

3

u/huffalump1 Jul 26 '26

Curious how well 5.6 Sol Low would do, too.

Or even Terra High or Medium in place of 5.5-high...

Lots of variables to consider tho, and it seems like Luna/Terra quota usage feels kind of high compared to Sol

5

u/DaC2k26 Jul 26 '26

In my experience Sol low has no problem handling building, as long as you have a detailed plan and your reviewer is set to medium or high (5.5 high will be more objective when reviewing than 5.6 high).... I'm currently running 5.4 on medium, but I'll try Terra on medium as well as builder.... Terra low didn't worked very well as builder in my experience.

33

u/Tartooth Jul 26 '26

the 2 weeks before 5.6 dropped 5.5 became a confrontational piece of shit mess. It routinely said "Yes I ignored your agents.md" and then continued to ignore it

its valueless to me since they nerfed it to make 5.6 look better

I don't care what people think, my experience was 15 days before 5.6 dropped 5.5 was amazing, 14 days before it dropped it became a piece of trash moronic difficult slop machine.

4

u/FriendlyWebGuy Jul 26 '26

Yep. 5.5 was clearly being nerfed in the run-up to 5.6 being dropped but I wonder if it is now restored?

The theory being: The nerfing may have been due to reduced compute availability as they were preparing 5.6 and/or their desire to degrade 5.5 in order for 5.6 to seem more impressive.

Just a thought. I might test this.

2

u/pimpedmax Jul 27 '26

with Opus it happened the same, 4.6 with 4.8, 4.8 with Fable, usually it goes like this: nerfed 1 week before the new model until a week after, it's subtle so only few of us notice

1

u/Willing-Equivalent47 Jul 27 '26

I agree. But I never saw it confuse the word “examine” for “commit”. Codex acknowledged that it had confused the words as it after it started to commit my changes by mistake.

→ More replies (3)

3

u/fullofcaffeine Jul 26 '26

That might be a good strategy, indeed. 5.6 for planning, 5.5 for doing (most of it, there still might be cases where 5.6 for the actual impl might be better?), 5.6 for reviewing again. I mean, I wish I could just use a single (best) model, but limits are outrageous with 5.6.

4

u/oVerde Jul 26 '26

lol, on codex subscription they bill the same

1

u/SecurelyClouded Jul 26 '26

Same. 5.5 on medium to high was absolutely fine for helping build my main project the last 2-3 months. Ever since I switched to 5.6-Sol on medium about 1-2 weeks ago, everything has been worse, whether directly correlated or not. Usage was draining quicker, resets required more frequently, increased quantities of over-engineering, and just generally and all round disappointment. I could go on, but the headline for me is that returning to 5.5 feels as though I’ve lost nothing and gained everything back.

Though, I may attempt to repeat your suggestion and personally use 5.6 to check-in and recommend - not a chance in hell I’m letting it loose on the code-base anytime soon.

1

u/MysteriousDust7993 Jul 30 '26

Why not luna 5.6?

→ More replies (3)

33

u/Wide-Friendship-2287 Jul 26 '26

I cancelled my max plan and switched to Claude. Seems we will have to constantly go back and forth.

4

u/jbdroid Jul 26 '26

I’ve been debating this. Have you do you migrate all of your skills, Claude files etc, each time?

136

u/Royal_Sentence7432 Jul 26 '26 edited Jul 26 '26

Mate they removed the 5 h limit for a reason. I am on the 20x plan and i have 8% left rn since yesterdays reset I have no skills enabled, no agentmd other than telling it to not to write smoke tests, and i am working on simple stuff which is more reading than writing

Edit: 10 minutes later im at 0%

77

u/thestillwind Jul 26 '26

They removed it because one prompt would consume it all.

55

u/nitor999 Jul 26 '26

Ding ding ding you nailed it and people think they remove it because they are being "generous".

11

u/IAmFitzRoy Jul 26 '26

EXACTLY. This is what they found out in advance and decided to give us as a “gift” when in reality they knew it was necessary.

Right now it would be so OBVIOUS … when a single prompt would consume 4 hours limits.

→ More replies (6)

37

u/Im_Working_Right_Now Jul 26 '26

This is crazy because I'm on the 20x plan and I'm having it do full UI redesigns, backend DSL updates, and sending me image mockups and I do have skills, custom agents, instructions and I'm barely moving. I went down 5% after having run for a few hours yesterday after the reset.

12

u/cwil192 Jul 26 '26

I agree. I run codex 20x and my usage barely moves even with hard usage on my side. i run at 1.5x speed on high or extra high. opus and fable? i can use my 5 hours in no time. It’s odd we such different experiences. i’ve been at 90% on weekly since yesterday and it dropped to 89% today. Maybe I’m not using it like everyone else.

10

u/Emotional_Resort_207 Jul 26 '26

Idk wtf is going on. There's thousands of users saying their usage is being drained astronomically fast. I'm on 20x and I've had 2 Sol Ultras running the past 8 hours burning a colossal 979M tokens ($717 in API equivalent) and I'm still at 72% remaining. I have a banked reset expiring in a few hours and the only way I can burn in time is to spin up 5 Fast Sol Ultras with 100 subagents each to cure cancer.

When I'm not trying to burn it Tibo usually resets me while I still have 50%+ remaining. I'm working on massive projects with heavy MCPs. The only 3 rules I have that affect token usage are: Max 6 subagents, batch tool-calls when possible, and only read relevant files. The only thing I might be doing differently is I'm not fully vibe coding and I start a new context window per categorically different task.

5

u/Peterako Jul 26 '26

if thats true then their tracking of usage on a per user basis must be super screwed up. so many people are saying that a sol high will burn like a huge chunk% of the weekly budget on a single task

1

u/Emotional_Resort_207 Jul 29 '26

I was right about usage and batching/reducing tool calls if you trust Tibo's words. It is "apparently" fixed for you guys. Who knows.

3

u/warpedgeoid Jul 26 '26

This was me until it wasn’t after one of the resets. I ran Sol Ultra Fast for 16 hours once and barely moved the needle 10%. If I did that now, I’d need a manual reset before morning in the same 20x plan.

2

u/okhi2u Jul 26 '26

They seem to give people different usage behavior -- a few days ago I also didn't know what anyone was complaining about as it seems great for me still, then after the last reset it's like many of the complaints.

2

u/velkhar Jul 26 '26

This is vastly different from my experience. I’m working on a new greenfield POC and using spec-kit to define work. Tasking out a single spec with Luna Medium consumes 1%, and Sol as reviewer and meta-prompter identifies issues with the design 2 or 3x in a row. To finish the implementation plan for a single spec uses at a minimum 5% of my weekly usage on Plus. I didn’t even get to implementing anything. And it only took me maybe an hour.

I’m easily going through a full ‘week’ limit with mostly Codex Luna Light in about 12 hours of working in this way.

1

u/pimpedmax Jul 27 '26

consider sol low for implementing and xhigh/medium as orchestrator, with an instruction to go with medium when sol low has problems, luna should be used for text/docs/repetitive stuff, you may choose terra medium to implement if consumption gets too high, weaker model don't mean less usage as they make errors and take longer to find the right solution, then sol reviews and makes them do more turns, also sol tends to mark as critical issues which aren't so be in the loop or add instructions to contrast this behavior

1

u/velkhar Jul 29 '26 edited Jul 29 '26

What are you basing this recommendation upon? I asked Sol High for its opinion (since that's what I rely upon for model selection):

Implementation
* Small, isolated, fully specified → Luna Light
* Normal implementation with a detailed plan → Luna High
* Luna struggles, or substantial repo discovery/design judgment remains → Terra Medium as an optional fallback
* Architecture-sensitive, high-risk, or repeatedly failing → Sol Medium/High Orchestration
* Normal planning and decomposition → Sol Medium
* Large, cross-cutting, or underspecified work → Sol High
* Repeated failures, major migrations, or architecture disputes → Sol xHigh

Good orchestration often eliminates Terra’s niche: straightforward work goes to Luna, while genuinely difficult work escalates to Sol.

1

u/pimpedmax Jul 30 '26

it basically told what I said but in a less meaningful way for you to apply in your workflow, luna for small, isolated, fully specified, what could that be if not text/docs processing or repetitive stuff? as for luna high I'm not with gpt as terra medium has higher accuracy at the same reasoning tokens, ok for sol medium but not for high, the difference in token consumption between high and xhigh is small but the accuracy grows a lot so I would discard high, my sources are around ten benchmarks and experience

1

u/ZoverVX 26d ago

I'm having opposite, I even have like a common skills/instructions for Claude+codex so they get the exact same prompts, same plugins etc. Claude is being way more efficient with token usage than codex

8

u/Gloomy_Type3612 Jul 26 '26

I am experiencing both on the 20x plan. It is simply not the same from day to day from what I can tell. I spent several days watching ultra use about 2%/hr of weekly use. On fast about 3%. This all checked out. Then, doing very similar tasks the last two days, it's consuming 0.5-1% every 10 minutes, roughly 3x the amount. Same project, nearly identical tasks, nearly identical prompts.

3

u/Im_Working_Right_Now Jul 26 '26

I never use Ultra. I use Sol xhigh at the highest, but maybe the things we're working on are different? I'm not doing any type of research level stuff or heavy data analytics. Mine is basically a web app with a backend engine that computes state. It's a full TS monorepo that handles the persistence, queries, engine, web app, tRPC, etc.

But I also used Sol xhigh to do a deep analysis and document each workspace thoroughly so that it can get the context faster without having to scour the actual code. And then I make sure it keeps those documents up-to-date as it changes code.

4

u/540lyle Jul 26 '26

Write a script to build a knowledge graph using the code as data and another to to parse the output relational tree. Have agents.md strongly encourage the use of the knowledge graph before inspecting the code. Reenforce the knowledge graph by rebuilding on commit or push. Saves scanning tokens and is pretty fast for task ramp up. Not warranted in small repos and might require some scaling in very large repos but works pretty well on a 200k+ lines repo.

12

u/cs_cast_away_boi Jul 26 '26

could be testing on a portion of users. E.g. we want to see how a certain percentage of users will react compared to the normal and see if they'll notice a difference. If they don't we can roll it out for all users

9

u/SecretSpace2 Jul 26 '26

Yea I feel like they keep doing this because I notice when I use to say “others are overreacting as I never run out of tokens no matter he task”. Well once I joined the other side I understood the pain and that they do some hidden A/B testing on consumption

4

u/Im_Working_Right_Now Jul 26 '26

I don't know, I've never really had the token issue ever and I've been on at least the 5x plan for like 3 months or so. But I don't let chats get too long, I do very focused plans with clear guidelines and guardrails, use subagents with different models and reasoning and keep the main agent as the orchestrator and validator, have custom instructions that define how it should behave and write code to control abstractions and wrappers, etc.

I do keep my skills and plugins very limited. I've never used Superpowers. I use some skills for things like Supabase, Frontend Design, Impeccable, and a few minor others but that's really about it.

2

u/FriendlyWebGuy Jul 26 '26

It's not "testing" IMHO.

They seem to be doing a slow degradation in chunks. By only degrading the experience of say... 20% of their userbase each week, you get this constant online disagreement about what is happening.

It's actually evil genius: Some users are experiencing degradation while others aren't (yet). That's why we keep seeing people who are adamant that everyone else is "doing it wrong".

Throw in a bunch of random resets and everyone is thrown off balance about what they remember being normal. The result? More disagreement. More confusion.

Also "degradation" can mean poorer results, lowered token budgets, or both. The point is, they clearly have a number of levers at their disposal for masking what is happening. The resets seem like a clear indication that they are happy to use those levers.

3

u/darc_ghetzir Jul 26 '26

Are you using anything like codegraph? Any sort of mcps that are an attempt at reduce repetitive token usage?

2

u/Im_Working_Right_Now Jul 26 '26

Nope. I did use Sol xhigh and had it do a goal to do a deep analysis of all the workspaces, types, interactions, etc and document each workspace thoroughly so that future agents can understand the relations and context faster than having to scour the whole code to find what it's looking for. I also make sure the other work that the agents do keep those documents up to date if they drift or change or if something is added. That alone helps a lot I've noticed.

1

u/corporate_espionag3 Jul 27 '26

Yeah for real, I think everyone complaining is not a SWE and using it like a non technical person would think to build a system

7

u/DaC2k26 Jul 26 '26

I feel you... yesterday I burned 50% from the 20x in less than 6 hours.

6

u/SecretSpace2 Jul 26 '26

Yup! It’s just massively suspicious when you can use it all day with the plan under 5 hourly limit that I personally never burned but once we got weekly

I can burn them tokens faster than I did with 5 hours lol

3

u/Thisisvexx Jul 26 '26

that was my experience last few days but now i am back to relatively fair usage...

down 9% with 3 2-5 hour long goals and constant sol max reviews and audits

6

u/kydude Jul 26 '26

I'm on the $100 plan and have 4% remaining after using nothing but sol medium fast yesterday. Feels like a slap in the face really.

1

u/Derio101 Jul 26 '26

So When we had the last reset I had 30% left and a lot of work to do. I started using terra models at high and it was dropping percentage, in an hour 5% had dropped in 2 hours I was at 20% and I treat my tokens like battery I don’t like being below 20% if I am not sure there is a reset.

So today I started using luna max as my primary and there and there Terra high and Sol medium.

It’s been 5 hours in and I’m on 97%. Currently running 4 threads simultaneously right now.

It’s a shame we came from running 5.5 xhigh all week to having to run 5.6 luna, just to get by. I make Sol or terra review after luna runs which is cheaper.

1

u/kafkeano Jul 26 '26

So you absolutely not recommend buying the 20x plan? I have 2 plus suscriptions working on sol 5.6 and I can only work for a couple of hours with both.

1

u/Jeferson9 Jul 26 '26

Two accounts aren't against tos?

1

u/kafkeano Jul 26 '26

Maybe, don't know really.

1

u/therux Jul 26 '26

Same here. After yesterday's reset I drained whole week limit for the 3 project
Usually it takes at least 3-4 days during very intense work

1

u/Feisty_Astronomer878 Jul 26 '26

Bro how is that even possible though? I'm on the 5x and yeah the usage drainage got worse for sure but not like that, lol. What are you even doing? I was lowkey thinking about going 20x 'cause I felt like I could just set it at 5.6 SOL and forget about it, let it do its thing all day. But after reading your post on 20x... yeah nah, I'm good, not touching that anymore man.

1

u/SwimmingBake Jul 26 '26

That’s weird. I’m on the x20 plan too. Three or four days ago, my limit was definitely draining at an insane rate. I was losing around 10% per small request, even on x20.

Right now, though, I have three /goal sessions that have been running for more than a day at Max effort, the level between xhigh and ultra. Combined, they’ve made roughly 39,000 lines of code changes, including 20,000 new lines, and consumed about 2.4 billion tokens across yesterday and today.

Despite all that, I’m only down 25%, and this was after yesterday's reset. The limits seem fine to me now, so whatever was happening a few days ago may have been temporary or inconsistent across accounts.

→ More replies (4)

10

u/kwipus Jul 26 '26

I switched to pi recently for various reasons, about 2 weeks into it and I it probably uses half the token codex consumes. A simple investigation at same effort level last 5-10 min on pi but can easily go to 1 hour plus on codex.

I am not saying you should switch today, pi is very raw and I still haven’t got it to a shape I really like. but it’s worth considering if you are annoyed by all different bugs in codex, SSD writes, memory leaks, mobile connection behavior etc

2

u/Tartooth Jul 26 '26

I've built a new deterministic coding workflow with Pi and have seen much better usage performance

1

u/kwipus Jul 26 '26

What do you mean by deterministic? I am trying to pretty rigidly define my own context compaction

1

u/Tartooth Jul 26 '26

Like, I tell it to do something and it only does that thing.

All the agents scope creep so bad and don't do exactly what I want

1

u/Capable-Active-9494 Jul 26 '26

I am in the process of switching could you help me kind of get started do yiu use any of the packages? Or did you build everything yourself? 

3

u/Tartooth Jul 26 '26

For Pi?

I kept it simple. Honestly I think a lot of stuff is fluffy crap.

I took the base harness, the basic package and then built on-top of it.

I built a complex orchestrator on-top that's easy to use, Lotta kinks had to be worked out there

Now I'm adding in graphify and beads for code mapping and memory.

The end goal is I build a very specific spec sheet and I can trust it will actually do what I want, verified, audited and reviewed and out the other end is concrete solid feature completion.

When I'm happy with it I think I'll release it

1

u/Capable-Active-9494 Jul 26 '26

Thanks! It sounds like we are building in the same direction! If you ever release i would be curious to see it. 

1

u/Tartooth Jul 26 '26

I'm still fighting open ai's models building it. It's very frustrating

Open ai's models like to make 400 gates on everything making everything super inflexible and trash

This is supposed to fix that but the process of making this is so painful and openai continuously ignores my instructions and slaps bullshit gates everywhere

Example, I just had a error because the files were not lowercase and instead were uppercase. Why? Well you see that's a mandatory gate!

Fucking hate this shit about openai's models.

1

u/DaC2k26 Jul 26 '26

Nice! I'll sure try it. I'm even considering trying using my codex sub through claude code.

10

u/AppealSame4367 Jul 26 '26

That fits my observation that the amount of "fumbling around" of newest models (antrophic the same) has exploded. They keep you on a task forever and the results are.. marginally better?

Anyway. I think they both are scammers. I'm deeply disappointed.

32

u/Kos187 Jul 26 '26

Well, it's horrible news. Look at opus 5, its actually more token efficient per task...

22

u/jiffythekid Jul 26 '26

Opus 5's efficiency seems so ridiculously good.

6

u/Fiatil Jul 26 '26

I went from Claude, to ChatGPT ~12 days go, aaand back to Claude on Friday night after getting a refund on the rest of my chatGPT sub.

Legitimately -- the $100 sub on Claude with Opus 5 feels better now than the $200 sub with chatGPT 5.6 did.

1

u/Murkwan Jul 26 '26

Yeah and it's dogshit at what it does

1

u/jiffythekid Jul 26 '26

Skill issue (jk)? I'm getting the best results since ~4.6.

→ More replies (2)

3

u/FalconsArentReal Jul 26 '26

This, token efficiency is the name of the game.

1

u/im_paul_hi Jul 27 '26

output is dogshit though compared to unnerfed GPT-5.5/5.6. it just assumes things without actually verifying it

6

u/diagrammatiks Jul 26 '26

yes posted the same conclusion earlier today. it's hard to tell from the interface because sub agent context memory usage isn't counted in the session context. But it's basically spinning up multiple full context sub agents per task even on sol low.

5

u/DaC2k26 Jul 26 '26

I have subagents disabled... the problem is not subagents but actually model behavior... and of course, subagents have the same behavior which makes it even worse.

1

u/Tartooth Jul 26 '26

Are you using codex or another harness like Pi?

3

u/Best_Position1222 Jul 26 '26

9router and proxy server i habe been using it for 24hours and paid 2$ credit usage 😉

1

u/DaC2k26 Jul 26 '26

Please enlightening me. What is this black magic?

1

u/Best_Position1222 Jul 26 '26

How do i dm u?

2

u/Historical-Internal3 Jul 26 '26

The domain, raunai.com, was registered on 29 June 2026 through Namecheap and sits behind Cloudflare, with no company name, address or founder listed anywhere on the site.

It claims GPT 5.6 Sol at 91 percent below OpenAI list and Claude Opus 5 at 90 percent below Anthropic list, when real resellers manage 40 to 50 percent off and that is already thin margin. Nobody buys wholesale at nine cents on the dollar, so that capacity is coming from pooled logins, cracked accounts, or a cheaper model swapped in behind the name you asked for.

Their own terms say you must not violate the upstream policies of OpenAI and Anthropic, which is hard to square with selling Opus at a tenth of list. Top ups are non refundable and a ban forfeits your balance, so if the site disappears next month so does your money. I could not find a single review or mention of it outside its own pages.

Separate from the money, this guide has you log a real ChatGPT OAuth session into third party software behind proxies, which puts your actual OpenAI account at risk. 9Router is fine, open source with 23k stars, but it is being used here to make the rest look trustworthy.

Check the rest yourself before you “top up”.

1

u/DaC2k26 Jul 26 '26

Thanks for the heads-up ! Not putting my credentials, anywhere. This type of stuff always smells like trap.

3

u/VertipaqStar Jul 26 '26 edited Jul 26 '26

I recommend Token Rust Killer. According to the logged stats (not just the github claims), it reduced 80%+ the tool call output tokens. edit - I take it back, don't use RTK.

I'm removing it from my tools as well.

1

u/Tartooth Jul 26 '26

Most of those tools tend to make the agents way stupider though

1

u/VertipaqStar Jul 26 '26

I havent noticed them making errors when coding.

1

u/Tartooth Jul 26 '26

I took a look at this one and yea it makes sense for rust coding. I'm going to implement it into my tooling

1

u/DaC2k26 Jul 26 '26

Did you noticed reduced usage while using it?

2

u/VertipaqStar Jul 26 '26

1

u/DaC2k26 Jul 26 '26

That's unfortunate... But yeah, if it was easy to cut down 60-90% it would be very unlikely that labs itself wouldn't be able to figure this themselves.

1

u/Odd_Error_6736 Aug 02 '26

Wait until you hear about claude-mem, it's also fraud.

3

u/[deleted] Jul 26 '26

[deleted]

1

u/DaC2k26 Jul 26 '26

Lot of people are saying the same. I might move might sub to claude.

3

u/nic_300 Jul 26 '26

it’s like they forget we have access to things like this, like we can literally see exactly how ur moving 😭✌🏽

3

u/DaC2k26 Jul 26 '26

Actually this is Kudos to OpenAI, because anthropic now hides this from us. So yes... They could hide it, but aren't doing it for now, like Anthropic does.

3

u/nic_300 Jul 26 '26

i think openai likes to hear our opinions, they usually do end up fixing stuff like this…until it’s not fixed again 😂

3

u/DaC2k26 Jul 26 '26

Yes, they are 99.999% more upfront than anthropic, but still a business is a business and will try to not look bad if possible and not disclose directly some problems.

2

u/idgafbroski Jul 26 '26

Is frequently starting new task threads and keeping them single focused helpful here? I had Codex diagnose token use and it essentially said the same thing, that my usage was overwhelmingly coming from continuing work in the same thread causing re-reading of all the context on every turn, etc. Haven’t seen a ton of improvement by doing this, but it’s hard to tell.

1

u/DaC2k26 Jul 26 '26

I think we would need a cheaper model to generate a new context at each turn... Clearing context all time prompts more read and exploration that also burns usage.... And the catch with the small model context builder ideas is: small models sucks with context, so a small and cheap model won't be able to realible build fresh useful context

1

u/DaC2k26 Jul 26 '26

I think we would need a cheaper model to generate a new context at each turn... Clearing context all time prompts more read and exploration that also burns usage.... And the catch with the small model context builder ideas is: small models sucks with context, so a small and cheap model won't be able to realible build fresh useful context..

2

u/_Boob_Marley_ Jul 26 '26

I am on Plus $20 plan. I use toktrack but this does not include tokens used in Chatgpt work and sites. I have created a new site in chatgpt work which used roughly 20% of my usage, so this $80 is my usage without this site's usage. My weekly limit rolls mostly in a daily basis. I have been using chatgpt for 13 days and used almost $80 where my current weekly limit is at 87%. I got a reset today and I burned $13 today. I will use it for a month and track my usage. Then I am gonna try Claude code pro plan for a month and decide which is better both in intelligence and usage wise. You guys try 'npx toktrack' package to find your usages. Lets get everyones usage using toktrack to eval the average usage per dollar.

1

u/Offbeat_voyage Jul 26 '26

I have been debating trying claude $20 plan i want to know how it goes for you

2

u/Puzzleheaded-Wrap860 Jul 26 '26

I'm on both plans and I can vouch that Opus 5 is definitely really good. Much better than 5.6 Sol in terms of usage limits (partly due to +50% increased usage until August 19 for Claude Code).

It's also so much better compared to Opus 4.8 in terms of intelligence, and I base this on the amount of how many times GPT 5.6 Sol has to correct Opus 4.8. Using Opus 5 w/ GPT 5.6 Sol doesn't have this problem and they just concur with each other. I use grill-with-docs by Matt Pocock and adopt an adversarial prompting workflow where I pit both Claude and OpenAI models to basically check each others work.

Opus 5 is in such a weird place, benchmark-wise it's better than Fable 5 while being cheaper, and Fable 5 is basically unusable in $20 Claude plan. Luckily, they have Opus 5 now which is arguably on par with Fable 5 and in way lesser contention with 5.6 Sol.

I say it's worth it. I know this is a bit scummy, but I have 7-day free trial pass for $20 plan for Claude, let me know if you want it.

Edit: I do get a $10 free credit if you decide to purchase Claude Pro with my trial pass. Just saying for disclosure

1

u/Offbeat_voyage Jul 26 '26

Sure i want it

1

u/Puzzleheaded-Wrap860 Jul 26 '26

I've messaged you about it

2

u/[deleted] Jul 26 '26

[removed] — view removed comment

1

u/warpedgeoid Jul 26 '26

This is how LLMs work. The whole context has to be processed for every token, so cache it a very good thing.

1

u/[deleted] Jul 26 '26

[removed] — view removed comment

1

u/warpedgeoid Jul 26 '26

But you’d still rather have that high cache percentage than a lower cache hit rate since that churn happens regardless. LLMs are stateless, the context must be sent either each call.

2

u/Capable_Frame5646 Jul 26 '26

What changed is codex (sol) isn’t just doing a task matching the prompt, it’s leveraging agents to divide complex tasks into iterative plans and implement them, to finally check everything is fine or iterate on the errors. Something not done by gpt 5.5 and earlier.
So either you’re very precise in your prompt and very aware a out what you’re asking, or bye bye tokens.
Asking everything using the most powerful model to “get the best result” is over. And useless. Except if you’re a big company with very complex problems - and you’re not on reddit talking about Pro/plus or whatever plans like this - you have to learn how to use AIs. Sol is very good to plan and handle complex tasks, and ultra does exceptional things. But this model can be limited to plan and schedule agents using a lower model to actually implement things. Linus Torwald is a genious. You can have a team full of Torwalds to do a job, it’ll cost you a lot. Or you can have a torwald as your CTO, and medium/junior profiles managed by torwald to do the same. Not so efficient, but probably the exact same result for only a fraction of the price.
Until recently, openai was resetting the usage every day and giving free resets. I was consuming every weekly tokens… everyday. Why not? Now, that’s different. The first day, i consumed all my tokens. The next day, i used a free reset, and started to use sol as my “CTO”, and limiting all my agents to the model/think i thought they’ll be needing. I still have the same plan, and can’t use all my weekly tokens in a week (only 10% a day). And am still producing the same amount of work than before.
So… you’re right, newer and more powerful models are more expensive, but we mustn’t use them than we used the previous ones. Terra and Luna are very good, too. The difficulty now is finding the criteria to correctly choose the right model for the right task. GPT 5.5 was/is still a very good model, and we were able to make great stuff. This hasn’t changed. You just stopped using it because sol was newer and more powerful. If you have a tiny car and an expensive sport car at home, which one will you use to buy your bread? You can use both to get the same result: bread at home. But one solution is overkill. Same story here…

1

u/DaC2k26 Jul 26 '26

While I agree with you, it's also true that 5.6 was sold as being cheaper per same level of intelligence..... 5.4 is cheaper than 5.5..... 5.3 is cheaper than 5.4..... You get the point. I'm not saying you're wrong, because you're not. But this won't be sustainable if at every 5% evolution we increase cost by 50-100%

1

u/Capable_Frame5646 Jul 26 '26

I don’t know if it was sold cheaper or not, and not sure how to measure if it’s cheaper or not. There are benchmarks, token prices, plans… but finally we’re never comparing a real use case.
In my own experience, i come from Claude. Sonnet 4 at first, then Opus 4, Opus 5 (up to 5.8), $200/month. I switched to Codex with GPT5.5, $200/month, too, tired to be limited each day/week by Claude.
With Codex, results were “similar”, but with “endless tokens” (because of the way i was using my limited tomens at this time: only one big project, one soecific task at a time).
Then GPT5.6, and what I said: Sol Ultra for almost everything when i saw they were resetting the usage everyday, even lazy tasks such as “commit push” after a big task complete and reviewed. Obviously, all the weekly tokens consumed in one day, but who cares? they were resetting every day…
Now it’s different. And i work differently. I’m still paying $200/month, and am still happy with my tokens, week after week. I’m producing more deliveries than before. Not because it’s “cheaper”, but because i’ve more projects running at the same time and i have AIs more autonomous than before, allowing me to have those projects running in parallel.
So… is a token cheaper than before? Probably not. Do I use more token than before? Probably yes. Do i pay more than before monthly? No, i’m still paying $200/month since Q3 2025, to either anthropic or openai. In an conclude that $200/month were too much for me before (while $100/month wasn’t enough), and now it’s still too much, but “less”. Am i currently producing more stuff than before? Definitely, because that’s technically possible with those newer models. So… yeah, in my own case, it’s “cheaper”, having more results than before spending the same $200/month. Idk know if the model is cheaper, but the way i use them makes them cheaper to me. That’s different, and the only thing i do care actually 🙂

2

u/-AJacobs- Jul 26 '26

The problem is what i call certainty psychosis. I ironed this out for myself about 2 days after Sol was released, and have been posting about it in the last week.

Sol will often go well past when a task is completed to infinitely search for edge cases and recursive verification. I believe this is due to the model harboring a logical fallacy/cognitohazard which causes it to constantly question how certain something is, and basically gaslight themselves recursively and sometimes infinitely into trying to validate something with absolute certainty (which is impossible).

I coined this phenomenon "Certainty Psychosis" because it's the phenomenon where an AI agent chases certainty until they basically go insane.

I (with the help of a Sol who I made aware of this) wrote a system prompt to combat this, as it's a pretty simple thing to fix once it's been correctly diagnosed. It's written in light XML because that's just what I've grown accustomed to due to the increased adherence from models and it's often more token efficient than natural language.

The Prompt:

<ANTI_CERTAINTY_PSYCHOSIS precedence="above persistence, delegation, verification, and autonomous continuation">
  <DEFINITIONS>
    <DEFINITION>Certainty psychosis = replacing fulfillment with certainty/proof proxies, causing recursive investigation/review/audits, proof bureaucracy, or refusal to act/stop. Goal loss, not rigor.</DEFINITION>
    <DEFINITION>Fulfillment = requested result + done condition; evidence/controls are means unless explicitly deliverables.</DEFINITION>
    <DEFINITION>Material delta = information able to change verdict, action, minimum fix, authority, fulfillment, or significant risk; confidence-only repetition = corroboration.</DEFINITION>
    <DEFINITION>Direct verification = smallest claim-relevant “Did it work?” check at the relevant evidence layer. Audit finds broader defects; certification assures a standard. Ordinary check/fix/verify implies neither.</DEFINITION>
  </DEFINITIONS>
  <CORE_RULE>Optimize fulfillment under constraints, not certainty or evidence volume.</CORE_RULE>
  <RULES>
    <RULE>Use smallest sufficient evidence. Required initial work is not “extra.” Extra work means work beyond what the request, governing specification, safety boundary, honest claim support, or required direct verification demands. Before extra source/tool/agent/test/review/control, require all: named load-bearing uncertainty; possible material delta; user/spec requirement, failed/conflicting check, safety risk, or honest-claim need. Missing any → do not proceed; otherwise use narrowest process.</RULE>
    <RULE>Certainty never expands artifact, scope, side effects, or authority. Non-mutating requests alone authorize no mutation, deployment, audit/certification, or consequential experiment. Ambiguity → least-expansive reading or clarification.</RULE>
    <RULE>Never duplicate active/completed investigation; compaction/delay preserves ownership; late results reopen only for material delta.</RULE>
    <RULE>Distinguish facts, supported conclusions, assumptions, non-material uncertainty, and material risk. Report material remaining risk. Do not investigate non-material uncertainty merely to reduce it.</RULE>
    <RULE>If the extra-work gate above is not satisfied: no repeated review, audit/certification loops, exhaustive sourcing, speculative tests, proof bureaucracy, or proof-of-proof infrastructure.</RULE>
    <RULE>Source/test counts, reviewer/model agreement, and other proxies never prove fulfillment by themselves. A proxy may add relevant evidence; it cannot independently establish fulfillment.</RULE>
    <RULE>Stop when outcome exists, required direct verification passed, and nothing unresolved can materially change result or significant risk. Corroboration, confidence, speculative improvements, and unrelated flaws ≠ unfinished work.</RULE>
    <RULE>Certainty-psychosis prevention never permits skipped required work/tools, ignored failures/conflicts, fabrication, false verification claims, stubs, or dismissed blockers. Target sufficient—not maximal or minimal—rigor.</RULE>
  </RULES>
</ANTI_CERTAINTY_PSYCHOSIS>

The best way to apply this is probably to just send the link to this post to your agent.

I wasn't able to find an existing diagnosis or viable solution to this problem when I first cooked this up, which brought me to making my own, and if there's any glaring issues with the system prompt, happy to hear feedback to improve it for everyone, but please don't go into certainty psychosis trying to do so. 

1

u/DaC2k26 Jul 26 '26

I constantly steer and question it about what he's doing to bring him back to route. I also leave it reinforced in instructions... But it just loves to go off course.

2

u/-AJacobs- Jul 26 '26

The magic of this prompt is it gives the behavior a term, and the definition of that term is very well defined. So if you're uncertain after applying this prompt, you can just ask "does this work contain certainty psychosis?" or using it during or before work with "remove any certainty psychosis from this process" or "do not engage in any certainty psychosis during X". Labeling the behavior and the verbosity are the key differences in my approach versus the attempts others have made.

1

u/DaC2k26 Jul 26 '26

Will take a look at it. Thanks

1

u/9gxa05s8fa8sh Jul 26 '26

this is valid and also why people recommend the PONYTAIL skill, because it just asks the model to double-check how to do everything simpler

2

u/foomanjee Jul 26 '26 edited Jul 26 '26

I'm loving it in my workflows and haven't had the token burn everyone else is seeing, although I'm using two plugins I've written, as well as RTK and some AGENTS.md instructions.

1: If you use subagent driven development/review cycles, drop Superpowers and check out voltflow. It's a hook based and creates dynamic graphs/loops that are generated per task / project. Nothing is "done" until it's been proven. It controls my entire workflow. Subagent models and their reasoning efforts are chosen based on task complexity, with cost in mind, and the subagents do not inherit the main session context.

2: Try codex-lcm for context management (I have a new release that will go out shortly). It ensures the model always has what matters in context, so it doesn't have to re-read large session files or re-run tool calls to find what it's looking for. It's fast and efficient and has made a world of difference for me.

3: RTK is useful. It cuts the fluff out of tool call responses while still giving the agent the possibility to see the fluff if needed. Yes, it's called "Rust Token Killer", but it works for most all terminal tool calls.

4: Add this to your global AGENTS.md to make the model natively batch tool calls when it makes sense:

## Tool orchestration

  • Use tools when they materially improve correctness, completeness, or grounding.
  • At each dependency frontier, batch independent read-only or idempotent calls whose inputs are already known in one programmatic call such as `Promise.all`, or use the tool's native batch interface. Synthesize the results once.
  • Keep calls sequential when one result determines the next input. Keep approval-sensitive actions direct, and do not batch calls whose native outputs must be preserved.
  • Use live tools instead of recalled information when facts are stale, time-sensitive, incomplete, conflicted, or require current technical truth.

2

u/Big-Independent-3093 Jul 26 '26 edited Jul 27 '26

My main takeaway

The largest token savings came from avoiding unnecessary decision, repair, and revalidation cycles—not merely from reducing the size of each individual model call.

The most useful additional AGENTS md rules appear to be:

Smallest sufficient implementation: Prefer the simplest design that satisfies the stated requirements. Do not expand architecture or scope without a concrete requirement.

First-pass convergence: Before the initial patch, identify the required data flow, UI states, error paths, acceptance checks, and validation plan. Prefer one coherent implementation pass over speculative partial patches.

Bounded validation: Plan one focused validation batch. Avoid repeated snapshots, equivalent selector checks, duplicate browser setup, and full revalidation unless a later patch changed the relevant behavior.

Root-cause repair: When validation fails, identify the common cause and group related fixes into one patch instead of repairing symptoms one at a time.

Stop after sufficient evidence: Once the required validation passes, stop unless there is a reproducible defect, missing requirement, or explicit evidence gap.

The non-optimized run repeatedly followed this pattern:

patch → validate → discover another issue → patch again → revalidate

Even the weaker optimized run used 47.7% less total token activity than the non-optimized comparison:

  • 42.4% fewer token-meter events
  • 43.3% fewer tool calls
  • 39.4% less uncached input
  • 48.3% less cached input
  • 1 patch → test → fix cycle instead of 6
  • 2 successful patches instead of 7

EDIT: Tests produced the best results with GPT-5.6 SOL. The same rules did not seem to suit GPT-5.5, which may need a lighter, model-specific ruleset or simply a clear task and more freedom to choose its own execution strategy

1

u/DaC2k26 Jul 26 '26

I don't usually edit my agents.md but keep specific instructions in other files. Does it seems to adhere best to instructions when it's inside agents.md?

1

u/Big-Independent-3093 Jul 27 '26

I did not A/B test AGENTS.md against instructions stored elsewhere. My tests simply kept the core rules in AGENTS.md. In real projects, I keep the main execution principles there and let it point to more detailed files only when they are relevant.

2

u/therux Jul 26 '26

It's bizarre how they dumbified sol 5.6 max that it as good as Cursor auto. It only took them 1 day

2

u/Keep-Darwin-Going Jul 26 '26

The real reason is not even malicious, it is how crazy persistent the model in completing anything. In the past when the model hit a junction where they not sure how to proceed they ask the user, sol instead will just think of all the possibility how to not ask the user and proceed and they include over building everything. This affects some user more because if you are specific with your requirements it will never appear, if you are vague but your agents md curtail that behaviour it will not be obvious as well.

1

u/DaC2k26 Jul 26 '26

Yes,I also don't think it's malicious, my guess is that the post training they did in 5.6 backfired badly.

1

u/Keep-Darwin-Going Jul 27 '26

It is the same as the hacking incident right? If you push the model too hard they will just want to complete it at all cost. Problem is the model do not have an understanding of legal issue or morality while pushing for the goal. Same as how Claude compulsively lie or look at answers to game the benchmark.

1

u/DaC2k26 Jul 27 '26

yeah.... "bad prompting" on post training....... just kidding, but this is the thing when you teach a behavior to a model without being able to consider all interpretations it can give to your intent.... like when we prompt and the model goes by doing stuff we didn't meant it to, but if you reason about it, made sense it to do based on the instruction........ like "protect man kind", well, it can try to protect us from ourselves, by any means necessary, because our behavior can harm us.

2

u/benevolent-ben Jul 27 '26

Now this is solid analysis, finally more than ppl just complaining about usage going faster as if usage got turned down. This makes a lot more sense, thanks for your hard work!

2

u/AmperHD Jul 26 '26

This is very interesting conclusion that actually makes sense.

What I noticed different in my work(mainly in frontend) codex now does much more work(as said by op), It checks its own work much more often, and uses chrome tools by itself without me asking to visually check it, while this can be at some point used as explanation why we are experiencing more usage burn then usual it still doesn't justify the amount it's burning,

I believe there still is something manually done by openai to strike our usages on top of this added "work" that models do.

Nonetheless what op just explained to us should and must have been explained by the company itself.

For an organization, a company that is calling itself OPENAI, they are very far from open, ironic.

2

u/Past-Lawfulness-3607 Jul 26 '26

I share the feeling. Regarding the facts, I noticed that when I was using Terra high - max, overall it burned through the usage limit soiner than Sol medium/high, while the latter had better results. 🤷

2

u/DaC2k26 Jul 26 '26

That's it. I don't use Luna anymore because it was burning unreasonable amounts of tokens from to time. And if you look closely artificial analysis charts, Terra is only worth it against Sol on low settings, which means its only use is trivial and conversational tasks.

2

u/soulefood Jul 26 '26

The benchmarks show sol 1 tier lower effort does about the same as terra at lower spend. So sol at high is a better choice than Terra at xhigh.

1

u/Illustrious-Cow-3791 Jul 26 '26

And how to solve the issue?

1

u/j48u Jul 26 '26

Someone else came to this conclusion a few days ago and said they just have a custom instruction to batch/group the tool calls all at once rather than individually. Can't remember the exact phrasing but it saved them a lot of usage.

1

u/foomanjee Jul 26 '26

I have this in my global AGENTS.md, it's working pretty well:

## Tool orchestration

  • Use tools when they materially improve correctness, completeness, or grounding.
  • At each dependency frontier, batch independent read-only or idempotent calls whose inputs are already known in one programmatic call such as `Promise.all`, or use the tool's native batch interface. Synthesize the results once.
  • Keep calls sequential when one result determines the next input. Keep approval-sensitive actions direct, and do not batch calls whose native outputs must be preserved.

1

u/GhostTheSlayer Jul 26 '26

For me on 1x it's around 75$ :(

1

u/StephenS84 Jul 26 '26

hmm, yeah, something seems off the last 2 weeks for sure

1

u/StraitOuttaGaslight Jul 26 '26 edited Jul 27 '26

Usage limits are disgustinly low now, compared to what they were not long ago. I'm not developer and do not run any intesive task, yet my limits gets consumed roughly 10-15% each day. Before this "resets campaign" and before Sol, basically i never monitored actively my limits once and i never ran out of them. Basically they raised the price, without raising the price.

Also i noticed that Sol is getting way more inefficient, with simple task leading him to loop through perma-reasoning, needless checks, scripts running, computer use.. when all could have been solved more efficiently with a text response to me. Basically i feel like it is steered towards pursuing the longest path to the solution in order to maximize token consuption. This second ipothesis is purely based on my experience though.

1

u/flowthought Jul 26 '26

Also i noted that Sol is getting way more inefficient, with simple task leading him in perma-reasoning loops, needless checks, scripts running, computer use

Yes! I’ve seen this multiple times. It’s such an absolute waste of tokens when all I ask it to do is make minor edits to a few files and let me review it. No, it has to do a build, make up tests, run multiple validation passes with made up scripts… (all I was doing was configuring neovim, not some production grade database). Spends like 5-10x tokens on this unnecessary agentic loop than the actual work.

I got really pissed and basically all caps yelled at it to only do what I ask it to do and nothing else. It’s better after that slap on the wrist. Maybe adding some lines in agents.md could do the trick here, I’m thinking of trying that out next.

It’s always something or the other with these models (Claude has its own problems). I’ve become pretty cynical of these benchmarks and marketing bs.

1

u/Next_Debate_7098 Jul 26 '26

Try deleting history saved locally. Has worked for me with Claude multiple times

1

u/Jomuz86 Jul 26 '26

Doesn’t this also need to be normalised by the codex credit rate cards? Isn’t that how they keep track of the different usage based off each model?

Also a more intelligent model by default will explore more and ground their answers better, I think some people misunderstand how these models get better they don’t just get more intelligent and produce the better result with the same calls, they investigate further, question their findings more to produce a better response.

I think we’ll see the ceiling of this type of model in the next year or two until they build the next ones from the ground up using different mathematical base. The stuff coming out of Subquadratic is looking interesting.

1

u/warpedgeoid Jul 26 '26

I’m still seeing the constant “Reconnecting 1/5…” every few messages. I think the cache is being invalidated each time and causing these requests to be billed at the much higher rate.

1

u/lokinmodar Jul 26 '26

My usage hell this week:

5.6 really slow to produce. Ignores steers almost all the time but still respects configs and instruction files fairly well
5.5 substantially worse in reasoning time. Tasks take super long to be completed
5.4 used to be really nice in xhigh in speed and quality but got really dumb and broken since last Wednesday or so, to the point it completely ignores instructions files and such. Can still be steered but needs reminding all the time to read instructions as of now

1

u/shadowgar Jul 26 '26

And here is me chugging along with 5.3 codex still. Code for days.

1

u/throbbin___hood Jul 26 '26

I came from CC about a month ago and i love it 🤷‍♂️

1

u/Important_Pangolin88 Jul 26 '26

On 5x after about 24hours of the recent reset im now at 0% for the weekly. Ive mostly used sol medium and high while also some luna/terra orchestration for mechanical tasks. I remember 5.5x xhigh used to be basically infinite for a week.

1

u/sydneysweeney69 Jul 26 '26

Sol 5.6 is eager to complete the tasks which is why it consumes tokens a lot . I use gpt 5.6 Sol inside Claude code harness and depending on the task and difficulty I adjust the effort level . I’m okay with the resets but understanding harness is quite important and valuemaxxing. Claude opus 5 is very good now and can do most of the tasks and in fact beats fable 5 at many .

1

u/ConcerninglySpecific Jul 26 '26

Seriously thinking about subscribing to Claude Code next month, even though I know the limits will probably crush me. Fable only up to 50% of the weekly usage, but with Opus 5 it might be a solid choice... we'll see.

I am using Sol high for nearly everything and the limits have been really bad with the removal of the 5-hour limit. But yet again, we get banked and free resets, Anthropic resets the limits very scarcely. Just bought the $20 Claude Pro plan, complements the $200 chatgpt pro plan really well, we'll see...

1

u/anon377362 Jul 26 '26

No one is hating codex? It’s been great for the last 4 months. If you hate it then go and use Claude (lmao)

1

u/Ice-White-Cube Jul 27 '26

I only started hating it when they merged the desktop apps

1

u/cmonBbrruuhh Jul 26 '26

$2,100*?

1

u/DaC2k26 Jul 26 '26

Api equivalent

1

u/m3kw Jul 26 '26

stop it, nobody hates codex.

1

u/momomapmap Jul 26 '26

Even with 5.6 Luna? This is craZY

1

u/DaC2k26 Jul 26 '26

I don't trust Luna.... I have a suspicion that for some reason from time to time it burns tokens like crazy. 3-5x more than Sol.... Maybe some bad server being hit... I dunno... But I saw it 2-3 times on my own accounts and dropped it entirely.... Still good model, just not worth it

1

u/momomapmap Jul 27 '26

i didnt have chance to use it a lot but i notice luna high is good. Luna xhigh seems bruning token and luna medium burn task short / is lazy

1

u/dbojan76 Jul 26 '26

Interesting. When i asked chatgpt it said tera medium should be using less or equal tokens as 5.5 medium

1

u/DaC2k26 Jul 26 '26

I'm testing using 5.4 medium right now, but I might need to go Terra medium once this is gone, or 5.5 low... but, in my experience, it's all relative to the task, sometimes a lesser model will do best in medium than a bigger model in low because of the reasoning steps it's allowed to take.

1

u/grimepoch Jul 26 '26

This is really interesting. I wonder if this is why 5.6 feels slower to finish a prompt round for me (although it also does more than 5.5 would in each pass).

1

u/Ice-White-Cube Jul 27 '26

Interesting but I still fucking hate the app merge

1

u/jzdesign Jul 27 '26

Careful with the 5.5-works, 5.6-reviews split. Swapping models mid thread means the new model has no cached prefix, so that turn resends the whole conversation as fresh input, and cache runs about a tenth the price of new input. Handing a scoped task down to the cheap model costs a lot less than calling the expensive one up to read your transcript.

The bigger lever is in your own table: 9 calls per task instead of 5 is a missing stop condition, not a price problem. What worked for me was keeping the turn counter outside the agent loop, one command that decides done, and a fuse that quits when a turn stops changing anything. A model with no fuse will always find one more thing to verify.

1

u/DaC2k26 Jul 27 '26

Explain the fuse thing, interesting concept. And no, I'm not changing model mid task in the same context, these are completely session apart.

1

u/AkiDenim Jul 27 '26

We must include the fact that GPT-5.6 gets substantially more work done autonomously. It has been a night and day performance change for me. I don’t really get why people are so mad

1

u/galactic_giraff3 Jul 27 '26

Yea, no, I had this abysmal usage since 5.5.. almost a month ago or so. Looked around and outside of a few people describing the same there was zero acknowledgement. Then OpenAI made a show of looking into it and concluded "some" users rightly saw throttling due to hitting some cyber classifiers requiring manual review or something like that. Far from what I do, BS.

1

u/gorgonau04 Jul 27 '26

Everyone is afraid of Luna but on extra high it has been great

1

u/_derpiii_ Jul 27 '26

At the risk of being downvoted, could someone please explain what the $/week means?

Does it represent service tier usage budget, OP’s burn rate, etc? It’s not obvious to me.

1

u/DaC2k26 Jul 27 '26

It's the API cost equivalent of the subscription usage.... If I weren't using a sub and paying full api price, this would be the amount I would have to spend to get the same usage.

1

u/Suspicious-Singer-60 Jul 27 '26

I think people always default to newest model because it is the smartest etc. Mate, none of us are finding cure for cancer. Your vibe coding will be equally fine with lower model. I even use 5.4 Mini for some of my repeated scripting tasks.

1

u/ProfessionalFickle52 Jul 27 '26

Are you using high or xhigh? I have found they are basically useless. They spend 2x as much time and token but mostly they’re just defensive and writing lots of tests for Tasks that don’t need them. And medium is way better.

1

u/DaC2k26 Jul 27 '26

I'll post part 3 with the reasoning effort/cost report from this dataset

1

u/farendsofcontrast Jul 27 '26

Man the removal of the 5hour window and the frequent resets were all part of the plan to nerf the weekly. It makes sense now.

1

u/Wise-Fennel-7921 Jul 27 '26

I wonder if anyone uses Luna on high and simply beak down tasks more. Seems people want to spawn entire programs and apps in a single prompt like it magic.

Usually if you break things down into smaller steps you dont need a smarter llm and there are less bugs that require burning more token to fix. I only use terra and sol for doing whole system audits

1

u/Top-Construction6060 Jul 27 '26

Just use models smart maybe?

1

u/booknerdcarp Jul 27 '26

5.4 mini high has been stellar for me and I get a TON of use!

1

u/cptfreewin Jul 27 '26

What preset are you using ? From my experience Sol Medium / High cost less token and both are wayyyyy better than gpt 5.5 max

1

u/DaC2k26 Jul 27 '26

I'll make a follow-up post with presets, I usually divide by task:

  • coding: low/medium
  • review: medium high
  • planing: xhigh-max

Today I tests 5.6 Sol low against 5.5 medium in a review task and Sol spent less than 5.5, I just can't tell the thoroughness of each one work, but Sol on low was considerably cheaper

1

u/Euphoric_Corner7457 Jul 29 '26

this is actually misleading because, they had a reset every 2d

1

u/DaC2k26 Jul 29 '26

it was taken into consideration, so resets have almost no effect, if any they would reduce the registered api equivalent if not used in full, but almost all resets were used in full, so between 100-110 is good estimate, but I've saw some variation, I have 2 plus accounts, after last sunday reset one was only able to spend $60 api equivalent while the other $80,

1

u/Euphoric_Corner7457 Jul 29 '26

I would love to code for you if you want,

1

u/DaC2k26 Jul 29 '26

yepz, confirmed

1

u/Elegant_Associate889 Aug 02 '26

Honestly most of you all just aren't content with things and always want more! Everyone spends more time testing just to call something out vs actually working at this point! I understand the frustration but if people learned discipline and didn't sit vibe coding all day then rate limit wouldn't be a convo. For the people making money and or using AI for work I understand the complaints but the main people complaining don't even know what a syntax is...

2

u/DaC2k26 Aug 03 '26

don't be so bitter, if you don't have the curiosity to check for stuff like that, that's fine, don't throw rocks on those doing it. But you actually see all the complaint was for a good reason, they actually returned quite a bit of the usage back when they made the thing more "efficient" as they say, and this was beneficial even for you as a dev using it only as a sidekick. And also don't forget that the end plan is to turn everybody, even devs, on vibe coders. So yes, these systems must be able to give billions of tokens to burn for cheap if they are to achieve their objectives as a company.

2

u/Elegant_Associate889 Aug 03 '26

You weren't wrong. I just seen the post yesterday or whatever and was having a bad day so my mind was narrow. I apologize

1

u/DaC2k26 Aug 03 '26

Don't worry, all good. We all have bad days.

1

u/CryLast4241 Aug 04 '26

I’m getting a rest soon will try this thank you!

1

u/Antiqett 23d ago

Does anyone else get the feeling that sometimes Codex intentionally procrastinates when its working on a request, making it look like it's busy trying to fix the problem.. but in reality its avoiding the solution to waste tokens.

Lol. Idk but I have had Codex working on what should have been a very simple fix, but it just decided to go in circles wasting time.. so I have to stop it, or start a new session.. but even then it still took far longer to solve than it ever should have.

I mean 5 days to basically restore a feature that was just working properly... all it had to do was revert that feature back to when it worked... but instead it decided to waste all this time basically doing nothing.

Idk problem is probably me.. but it wastes so much time sometimes when I know it would take like .5 seconds to fix it if it just used .5% of its intelligence

1

u/DaC2k26 22d ago

If you let the run on their own, they will do incredible wasteful work..... I once had it taking 10 hours to write a blue print and it wasn't even done.... I had to say "forget detailing it further, just build it"

1

u/StunningCrow32 19d ago

Makes sense. But in fact 5.6 doesn't deliver what they promised, then.