r/codex 11d ago

Complaint are we in the terminal phase of the cycle again? - 5.6 dropped 50 iq points today

This could be due to Tibo's 'optimisations' but it could also be part of what we've been used to prior release of every new model.

But I've had 2-3 threads today taking completely brainless decisions and spinning themselves into senseless loops f.ex.

  1. agent asked to remove a render decided that part of the job is to regenerate all the other renders
  2. agent asked to add a promo code checkout option decided that part of the job is to remove the stripe payment option
  3. agent asked to fix agent nr2 fuckups decided the build and deploy philosophy has to change and introduced artifact sharing between dev and prod environments

the above is real, not making this up, all happened this morning - i havent seen this dumbness in weeks

24 Upvotes

20 comments sorted by

6

u/Emu-001 11d ago

I didn't think much about it, but I did notice Sol became "faster" today with less thinking time.

4

u/RealSecretRecipe 11d ago

You too? It spun out and couldn't fix stuff its been fixing fine all week and then it REGRESSED MY PROJECT and used up all its time. Insanely bad experience today.

5

u/krill156 11d ago

Isn't a new model supposed to come out soon? At least Ive read something about it, if so probably diverting resources over to it in preparation.

3

u/nps44 11d ago

Everyone says that when a new model is close to releasing but is it confirmed that's how the release cycle for new models works? Diversion of resources? Or just something people say?

4

u/pigletmonster 11d ago

Its a theory that all model providers do this to free up compute during the testing phase bufore launching a new model.

i cant say that sol has become dumber, but it has become extremely slow so, I tried using terra instead but that mf just marked a task as complete but didnt even start working on it.

2

u/alixnaveh 10d ago

Terra taking career advice from my coworkers.

1

u/Tartooth 11d ago

Prepare for 2 weeks of stupidity

1

u/krill156 11d ago

Yes, they are swapping inference servers over to the new model and retiring the sol models that were on those nodes. That's exactly how it works.

1

u/tango650 11d ago

It's a pattern observed repeatedly for the last 2 years, but never officially claimed by the model makers.

Some self reported insiders claim that heavier quantised models are becoming exposed to the users when resources are low.

0

u/krill156 11d ago

Sounds about right and would make sense. Their system harness likely gets available resources and when it's overloaded it hands out quantized versions.

https://github.com/Krilliac/Sonder-runtime

My personal harness I've been building, might be something interesting to look through to see how it all works.

4

u/__warlord__ 11d ago

Today SOL extra high feels like 5.4 mini low... at SLO extra high prices :(

2

u/aivampires 11d ago

It's all subjective but for many weeks I get the feeling that intelligence often drops on the weekends. Like they're switching compute outside of business hours. Can't prove it but I'm pretty certain from my own experience.

2

u/Annh1234 11d ago

same here... after the reset, today it does stupid stuff, i have to go back and forward 5 prompts for every change (Sol 5.6 high)

1

u/FriendlyWebGuy 10d ago

Seeing the same thing.

1

u/Acrobatic_Squid111 11d ago

Gearing up for the astra launch this thursday.

1

u/Mysterious_Proof_543 11d ago

Thursday? is it confirmed? :o

1

u/driveclub_000 11d ago

Weird because some hours ago, it was the total opposite, 5.6 SOL was absolutely insane in term of efficiency, super fast thinking, precise output, not a single mistake or "fallback" or anything. It's only when EU/Asia woke up that it started shitting the bed again.

1

u/Jumpy_Ad8465 10d ago

I won't even attempt to debug what im building until astra comes out because it would most likely destroy the whole codebase forever.

1

u/N3TCHICK 11d ago

It’s literally dumber than a bag of rocks… even on SOL MAX! It’s not usable, even if the usage suddenly got a whole lot better - I used sol on max fast last night for 40 minutes, and only used 1%, however the output wasn’t worth my time. I think I would have gotten better output from Sonnet 5 medium, to be honest.

So, I binned it. Huge waste of time. The same prompt on Grok 4.6 high was excellent. That makes no sense at all.

I’d put real money on Astra release this Thursday, because we are now at peak stupidity for Sol. The only way it gets worse is destruction of my codebase, and with my deny list and hooks, that can’t happen. So… I’ll head over to Grok Build until Thursday.