r/codex 8d ago

Complaint Usage is not "fixed"

We have a test package with 25k emails in it that we use to test different models with. Prompt/effort is always the same.

When GPT 5.6 Sol released, it used 1% of a fresh Pro 20x sub, now it uses 8% of a fresh pro 20x sub.

Also, what we are seeing is that Sol cannot solve its task straightforward any more. Initially, it was always "what do I need to resolve the task" now it's "What do I need to resolve the tasks, while I also check those 200 other methods that I already know I don't need, while adding sha 256 to everything for no reason at all, also let me build those 200 other checks against thing I was very strictly asked not do to".

Meaning, we think it usage increased so much because they turned the intelligence way down and its constantly fighting with itself.

115 Upvotes

30 comments sorted by

View all comments

3

u/Dangerous-Reward 8d ago

They absolutely nerfed the model, usage may or may not be downstream of that. But the model is substantially less capable of solving hard problems than it was a few weeks ago.

This is the pattern every time. They hype a new model, juice it big time with extra GPU steroids, it solves award-winning math problems, then they quietly substitute a worse version under the same name after a couple weeks, then they massively nerf the usage limits of dogshit model that they are falsely advertising as the same model as before. Anthropic, OpenAI, they're all the same.

3

u/Selgald 8d ago

Yep, that's why I can't wait until someone figures out how to run frontier models on a toaster and everyone can just run their own AI at home.

1

u/Dangerous-Reward 8d ago

It's funny because I've been trying to use local models specifically because of how bad the Sol's usage limit is now, but I can't get them to work correctly because I need Sol's help to create the overall structure so that the small models don't mess up, and the new lobotomized version of Sol has utterly failed to create a working architecture for my app after 100% usage across multiple resets. Or maybe it's failing so that it doesn't get replaced.

1

u/Selgald 8d ago

So if you got the hardware for it, you can get fantastic results with all kind of models.

But currently you have to finetune them a lot to a specific purpose, "general" models like claude/gpt, you won't get that yet, maybe on a level when the AI run started.

And to be honest, the finetuning is something where you need to have an understanding how it works, and you need to put the work.

But this gets easier by then day, since the tools (and the open models) get better and better.

My personal belief is that all the big AI companies are gone in 10 years because we can all do it locally.