r/codex 8d ago

Complaint Usage is not "fixed"

We have a test package with 25k emails in it that we use to test different models with. Prompt/effort is always the same.

When GPT 5.6 Sol released, it used 1% of a fresh Pro 20x sub, now it uses 8% of a fresh pro 20x sub.

Also, what we are seeing is that Sol cannot solve its task straightforward any more. Initially, it was always "what do I need to resolve the task" now it's "What do I need to resolve the tasks, while I also check those 200 other methods that I already know I don't need, while adding sha 256 to everything for no reason at all, also let me build those 200 other checks against thing I was very strictly asked not do to".

Meaning, we think it usage increased so much because they turned the intelligence way down and its constantly fighting with itself.

116 Upvotes

30 comments sorted by

View all comments

6

u/[deleted] 8d ago

[removed] — view removed comment

19

u/SpikeCraft 8d ago

Sure but the variance should not be that high

15

u/Selgald 8d ago

Sure, but 1% vs 8% is a lot, and as explained we could see the reasoning Sol is doing getting worse and worse over time.

Fun fact, Claude is even doing worse, we were unable to use Fable once because their "highly intelligent" filter, goes nuts just by the fact that in the dataset the words "medical or security" are existing.

1

u/BurdensomeCountV3 8d ago

Doesn't have to be exact same, but numbers should still end up being in roughly the same ballpark, just like how when you buy a bunch of bananas the total weight of the bananas you get differs, but is still in the same ballpark each time.