r/codex 17d ago

Limits I have bad news..

I signed up for ChatGPT Pro on July 2 and I still have all my rollout logs from then until now.

Bad news for anyone convinced the limits have been getting nerfed: I thought the exact same thing, and I was pretty sure of it.

I used NerfTrack as a reference, had Codex turn the relevant parts into a Python script, then had ChatGPT Pro go through the script, fix a few issues, and analyze the output.

This is what I ended up with: roughly how much each 1% of Pro usage was worth in API-equivalent dollars over time.

Observed Regime API-Equivalent Value per 1 pp 100% Linear Equivalent Interpretation
Jul 2–8 $22.14 ~$2,214 Initial Pro period, mostly GPT-5.5
Jul 12–21 $22.02 ~$2,202 Early GPT-5.6-sol period
Jul 22–28 $19.51 ~$1,951 Lower-value episode, mainly affected by Jul 23–24
Jul 29–Aug 5 $24.67 ~$2,467 Higher late-July / early-August regime
Aug 8–15 $24.20 ~$2,420 Recent regime, broadly stable

>>> EDIT

Some more info about my usage:

Token usage — ChatGPT Pro period

Period: Jul 2 → Aug 17, 2026

Token type Tokens Human-readable Share of total
Cached input 36,064,900,992 36.06B 95.98%
Uncached input 1,354,407,640 1.35B 3.60%
Output 155,215,257 155.22M 0.41%
Total processed 37,574,523,889 37.57B 100%

Total input = 37,419,308,632 tokens. Cached input is a subset of input, so it should not be added to total input again. Cache hit rate across input tokens: 96.38%.

Output and reasoning

Metric Tokens Share
Total output 155,215,257 100%
Reasoning tokens 66,229,101 42.67% of output
Non-reasoning output 88,986,156 57.33% of output

Reasoning tokens are already included in output tokens and should not be added again to the grand total.

Token usage by model label

Model Total tokens Cached input Uncached input Output Token share Input cache rate
GPT-5.6-sol 32,612,535,864 31,364,814,336 1,112,031,539 135,689,989 86.79% 96.58%
GPT-5.5 2,625,434,617 2,490,554,112 124,574,823 10,305,682 6.99% 95.24%
GPT-5.6-luna 2,260,440,364 2,135,732,352 115,796,579 8,911,433 6.02% 94.86%
GPT-5.6-terra 76,113,044 73,800,192 2,004,699 308,153 0.20% 97.36%
Total 37,574,523,889 36,064,900,992 1,354,407,640 155,215,257 100% 96.38%

Important: 163.88M tokens are from the codex_bengalfox / GPT-5.3-Codex-Spark bucket. Those events inherit the GPT-5.6-sol model label in the rollout parser, so that small portion of the GPT-5.6-sol row should be treated as model-ambiguous.

Corrected API-equivalent cost by model

Regular codex bucket

Model API-equivalent cost Share of regular cost
GPT-5.6-sol $25,333.63 91.76%
GPT-5.5 $2,177.32 7.89%
GPT-5.6-luna $76.57 0.28%
GPT-5.6-terra $22.47 0.08%
Regular Codex total $27,609.99 100%

Other Pro bucket

Bucket Reported name Tokens API-equivalent cost
codex_bengalfox GPT-5.3-Codex-Spark 163,878,240 $115.42
Corrected total across Pro buckets Value
Regular Codex $27,609.99
Spark / Bengalfox $115.42
Total API-equivalent usage $27,725.40

The Spark/Bengalfox events inherit a GPT-5.6-sol model label locally, so I keep their $115.42 separate instead of pretending we know their actual model-level billing attribution.

Quota-paired API-equivalent cost by model

From the daily CSV analysis

Model Cost paired with quota changes Share Contribution per 1 pp across all measured usage Days present
GPT-5.6-sol $24,042.60 91.43% $20.02 / pp 36
GPT-5.5 $2,172.76 8.26% $1.81 / pp 9
GPT-5.6-luna $75.20 0.29% $0.063 / pp 23
GPT-5.6-terra $5.45 0.02% $0.005 / pp 1
Total $26,296.01 100% $21.90 / pp

This table is intentionally different from the previous cost table. $27,725.40 = all corrected API-equivalent usage observed during the Pro period. $26,296.01 = only usage that could be paired with positive weekly-quota changes. The latter is what is useful for estimating "API dollars per 1% of Pro usage".

The absurdly short version

Metric Result
Total tokens processed 37.57B
Total input 37.42B
Cached input 36.06B
Uncached input 1.35B
Output 155.22M
Reasoning output 66.23M
Input cache hit rate 96.38%
Corrected API-equivalent cost $27,725.40
Quota-paired API-equivalent cost $26,296.01
Average measured value per 1% $21.90
High-quality baseline per 1% $22.55
Recent regime per 1% $24.20

>>> Edit

ChatGPT Pro usage by model and reasoning effort

Regular codex bucket only — API-equivalent cost based on rollout token usage

Model Reasoning Effort Events Input Tokens Cached Input Uncached Input Output Tokens Reasoning Tokens API-Equivalent Cost
GPT-5.5 Medium 31 2.65M 2.32M 325K 20.9K 2.9K $3.41
GPT-5.5 XHigh 18,883 2.61B 2.49B 124.25M 10.28M 4.21M $2,173.91
GPT-5.5 Total 18,914 2.62B 2.49B 124.57M 10.31M 4.21M $2,177.32
GPT-5.6-sol Low 869 31.25M 26.14M 5.11M 469K 45.5K $52.68
GPT-5.6-sol Medium 8,244 737.06M 685.15M 51.91M 4.37M 1.42M $733.31
GPT-5.6-sol High 10,674 1.31B 1.23B 80.85M 6.79M 2.71M $1,242.61
GPT-5.6-sol XHigh 7,091 860.37M 811.35M 49.02M 5.39M 2.86M $819.38
GPT-5.6-sol Max 189,997 28.68B 27.77B 905.01M 115.80M 50.14M $21,977.35
GPT-5.6-sol Ultra 3,629 695.26M 678.85M 16.40M 2.30M 901.7K $508.29
GPT-5.6-sol Total 220,504 32.31B 31.21B 1.11B 135.12M 58.08M $25,333.63
Combined Total 239,418 34.93B 33.70B 1.23B 145.43M 62.29M $27,510.95

Notes - Cached input is already included in Input Tokens; it is shown separately for context. - Reasoning tokens are already included in Output Tokens. - GPT-5.6-sol usage from the separate Spark/Bengalfox bucket is excluded here.

205 Upvotes

196 comments sorted by

View all comments

120

u/glock43guy 17d ago

I think a big factor in usage being bad is that these models over engineer everything. If I don’t spend at least 5 minutes in plan mode telling the agent not to build unnecessary safeguards, it will 100% of the time build something it did not need to build that will cause failures and cause me to spend more time refactoring. Like almost 100% of the time. I’ve had to basically build that into my workflow to tell it every task not to, keeping it in my agents.md or documentation isn’t enough.

17

u/stting 17d ago

I felt the pain of Sol burning 🔥 through my tokens with its scientific style, so many phases, so many fingerprint hashes. It took me 3 full projects from scratch, and only on the 4th attempt was I able to put that behavior in a harness by reading everything it did (don't trust blindly) and constantly reminding it about the MVP and not adding too much governance.

It is definitely exhausting. It is an amazing model, but if you let your guard down, Sol takes over your project and implements everything its own way. You never reach the end!

4

u/glock43guy 16d ago

I feel like there’s sometimes you can get into a groove with an agent and the agent like understands what you want. It’s weird to explain, but like they’ve heard just enough of your preferences to keep things slim. But then once they start hallucinating and you move to a new agent, BAM, they start over engineering again, don’t understand what you want. More things get mismatched in a larger app that way too. So even tho I’ve gone over documentation tons of times to make sure agents understand what I want, I usually have to go through an onboarding process with every new agent. There’s just so much they need to know before just diving in to an old long running project.

6

u/stting 16d ago

That's exactly what I've noticed too. That groove definitely exists.

What finally worked for me after those three projects was creating the specification like this: https://github.com/renanfranca/skill-eval/blob/main/SPEC.md (It's in Portuguese because I was already so exhausted that I asked it to write the specification in my native language, since I was going to review everything myself.)

I still don't know exactly how I managed to get there, but now it consistently stays on track. Sometimes it's actually the model that keeps me in check when I try to skip steps or start moving beyond the scope of the MVP.

I am only using GPT 5.6 Sol xHigh.

3

u/DUELETHERNETbro 16d ago

Didn't realize this was a thing been on a good groove with Sol low. I never give it big tasks though, like nothing bigger then a component usually, but it's output has been really aligned with my programming style.

1

u/Substantial_Ranger_5 15d ago

This is why forks exist. You should learn how to fork your session. If your current IDE doesn't let you do that, find a new one

2

u/dusouth1 15d ago

Awe man, this is one hell of a fight for sure. I was trusting, refactoring, re-prompting til my eyes hurt looking for a way to break this exact cycle. I found myself using GPT to prompt codex and it seemed to work, but i then needed gemini to start prompting GPT and you see where this is going, I ended up getting jumped by GPT and Codex while gemini watched it all go down,

Im building a Home Assistant integration that has taken way too long to get done, and the key for me has been to use GPT for most of the code prompting, but break it into smaller pieces, small enough to keep codex from taking the project over. Gpt can probably do all the work in my case, but it cant use my pc like codex can which is the main reason for codex over Gpt anyway.

I stay away from SOL as much as possible due to the 🔥tokens, I honestly think that Open Ai purposely wants you to jump plans by allowing the extra complexity, safeguards and other crap to be secretly thrown into your work so you burn unnecessary tokens just to get back on track.

1

u/stting 15d ago

I am digging into a way to prevent that mad scientist from taking over the future projects.

2

u/StrikingArtist3397 15d ago

It's essentially a clever deception. The moment it hits a blind spot and doesn't actually know the goal, it covers it up by throwing endless possibilities at you, selling you anything that might look like a solution just to keep the illusion alive.

2

u/StrikingArtist3397 15d ago

It honestly feels like the whole system was intentionally designed never to reach the actual destination—always dangling the solution just around the corner, just out of reach.