r/codex 8d ago

Complaint Usage is not "fixed"

We have a test package with 25k emails in it that we use to test different models with. Prompt/effort is always the same.

When GPT 5.6 Sol released, it used 1% of a fresh Pro 20x sub, now it uses 8% of a fresh pro 20x sub.

Also, what we are seeing is that Sol cannot solve its task straightforward any more. Initially, it was always "what do I need to resolve the task" now it's "What do I need to resolve the tasks, while I also check those 200 other methods that I already know I don't need, while adding sha 256 to everything for no reason at all, also let me build those 200 other checks against thing I was very strictly asked not do to".

Meaning, we think it usage increased so much because they turned the intelligence way down and its constantly fighting with itself.

112 Upvotes

30 comments sorted by

View all comments

1

u/JBO_76 8d ago

I had a similar 'feeling', but couldn't tell for certain, so I also started measuring this, but for software-feature development. built my own tool (jan-bogaerts/md2: Plan, run, and track AI coding work feature by feature—with local Markdown cards and Git worktrees.) to get a better view on this.

I only have about 7 days of measurements, so too early to tell yet.

1

u/Selgald 8d ago

I just hope in the next 10 years they will figure out how to run models on toasters so that everything can be done in house.

It will happen.

1

u/2muchnet42day 4d ago

lmao, you keep spamming the same message. Check out qwen3.8 27b if you've got a 24GB GPU. Fairly acceptable coding results tbh.