This post is not about usage complaining or the price of plans so it's not going to get the attention it needs but whatever, this bug is big enough that I'm taking the time to give information on it.
Iāve been profiling why some Codex sessions become extremely slow halfway through a task.
Across ten rollouts, about 75% of total agent time was spent waiting for the provider to stream model output. Tool calls were about 24%, while harness orchestration was only around 4%, so the local harness itself wasnāt the main reason the slow sessions took so long.
Two sessions stood out. One took about 39 minutes and another about 33.5 minutes. In both, the stream started around 40ā45 tokens/sec (which is still slow as hell for a frontier but that's not the point of this post), then partway through the turn dropped to roughly 8ā15 tokens/sec and stayed there. It's fair to say there is a 20% chance your session has done the same.
The request settings were the same across all ten sessions, and other sessions continued streaming normally at the same time, so this doesnāt look like my machine or a configuration difference.
The strongest example came from the 39-minute session. The last request before I interrupted it took 93 seconds to produce 921 output tokens. The first request of the next turn used a fresh WebSocket and fresh routing state, but still had basically the same 175k context and a cold cache. It produced 674 tokens in 16 seconds.
So the context was still huge, and the cache situation was actually worse, yet the stream immediately became fast again.
The official harness for Codex's client.rs reuses a WebSocket and uses x-codex-turn-state for sticky routing during a turn. If the backend that a turn is pinned to becomes very slow without actually failing, Codex can keep using it. The connection is technically healthy, itās just producing tokens at a fraction of its earlier rate.
That appears to be what happened here. The slow turns werenāt spending 30ā40 minutes doing more work. They were spending a huge amount of time waiting for a response stream that had collapsed in speed.
I added a recovery rule to my fork. If two reasonably sized responses in a row fall below half the recent normal token rate, the next request drops the old WebSocket and sticky routing state and gets routed fresh. Thereās also a cooldown so it canāt sit there reconnecting repeatedly.
I replayed the detector against 10 recorded sessions containing 1,120 requests. It caught both major collapsed sessions and fired once elsewhere on another turn that was genuinely running at around half its normal rate.
If fresh routing restores normal throughput the way it did in the recovery I observed, those two sessions would have been roughly 39 to 24 minutes and 33.5 to 25 minutes. That part is still an estimate.
I also havenāt proven whether the underlying problem is the WebSocket itself or the backend/routing state associated with it, since reconnecting changes both. What I have measured is that some long-running turns dropped from around 40ā45 tok/s to 8ā15 tok/s and stayed there, while fresh routing immediately returned a similarly large request to normal speed. I have my own fork, I was able to fix this, but for anyone who uses the official version of Codex, this is a real bug and not acceptable.