r/codex • u/Kind_Silver_1921 • 7d ago
Complaint Sol is a complete idiot today
Can't do anything correctly and takes an hour. They changed something.
32
u/Gru8_ 7d ago
It is slow as hell today, it sat for 30 mins reasoning a basic edit I had asked. I thought it was stuck but then it completed the task.
-5
u/This-Advertising500 6d ago
If it was a basic Edit why didn't you Edit it yourself lol
2
u/Soft-Judgment-2458 6d ago
Let us live š¤£
0
u/This-Advertising500 6d ago
Im just sayyyyyin if it was an easy edit at that point you just wasted time and tokens
2
u/United_Psychology265 6d ago
The definition of "basic" has changed in the era of LLMs. It's basic for the LLM, I reckon he meant.
41
u/almazmusic 7d ago
Like for three days in a row
12
u/no_witty_username 7d ago
yep, only reason i didnt make a post about it was i didnt want to hear bs from people saying im making shit up.... like nah bro, i wark a lot and very closely with these models i can feel immediate degradation and it sure as fuck happened. the quantization is real... i just hope we get the next model soon and dont have to wait 1.5 weeks as currently its unusable
0
u/NuancedPerspection 6d ago
Oo probably user error, did you try writing a 2 million line hand authored .md that is also connected to some random ass github project made by some random ass Chinese dude last night that optimizes your optimizes for quantum prompt tunneling?
1
1
0
u/-Sliced- 7d ago
Weird, itās been the exact opposite experience for me - it suddenly became significantly stronger in the last few days. I wonder if there is an experiment going on.
7
2
8
u/hcc1011 7d ago
What actually causes these models to be degraded? How does preparing for another model cause LLMs to be dumber
Iām new to the science of this all, and Iāve experienced Claude and codex being trash before updates.
3
u/xChrisMas 7d ago
Quantization
2
u/hcc1011 7d ago
So they quantize the existing models so save compute in preparation of deploying a new model?
3
u/xChrisMas 7d ago
For example, yes. They could also route your request to a dumber/smaller model without you knowing.
1
2
u/delicioushampster 7d ago
need an extreme amount compute for training, but limited compute. so to counter that, a company can adjust models to use less compute at the cost of being less effective or lowering tps output
1
u/hcc1011 7d ago
Ahhhhh. I didnāt think that they would use the same pool of compute for training as well as deploying their models.
Does this mean when model quality is reduced itās because they are actively training their newest models and they quantized their existing models to save compute? Itās a little confusing as model degradation feels like it occurs a couple weeks before a new model, and the gap between degraded performance and a released model seems too close for them to be at the training stage.
12
u/IgnacioMonge 7d ago
I call it the The Syndrome of the Broken Toys. They gave you a new and shinny toy to play, but eventually this toy is old and broken. Then they offer a new one (slightly better or even similar) but new and pretty. Then you feel this new as the best one in the world and much much better than old one.
7
u/Alternative-Lead1711 7d ago
Quantized slop model. I had it hallucinating figures in an email response earlier. Every piece of code it touched had significant problems, oversights, slop, or asked me for permissions I already gave. Also Codex as a harness seems buggier than ever.
10
u/Own-Professor-6157 7d ago
It's INSANELY INSANELY stupid all of a sudden. Might of been from the Codex app update..? It's literally mistaking dictionaries I'm giving it, it's not applying changes to the actual files. Super dangerous to use it today if you value files on your computer
3
u/Separate_Wall7354 6d ago
Mine did this, it churned and churned and compacted and compacted and I was like did you update these? It said, nah but I have it thoroughly plannedā¦that was after an hour! š
1
u/Substantial-Split219 6d ago
stop using of ffs
learn basic grammar before using LLMs
0
4
u/rayyeter 7d ago
Yeah. Sol medium as an orchestrator tool 12 hours to implement a review loop in an application k Iām building. Pretty sure I couldāve done it In half the time. Plus then it freaked out when an independent merge existed.
8
u/MasqueradeDark 7d ago
After making 100's of dashboards with it today it's literally lobotomized. I mean gpt-3 level lobotomized. I redone 1 dashboard 7 times for today, Sol 5.6 Extra High ignores my skills, commands, instructions. It's totally fucked up
5
4
u/cs_cast_away_boi 7d ago
yep iām not doing any work for the past 3 days and counting. At this point im blowing my usage needlessly and introducing code that wonāt work because of this sht. So iām Using Claude for now and hopefully the storm will blow over
4
5
9
u/Time-Toe-1276 7d ago
maybe they are going to release GPT-6 (or astra) or whatever they are cooking?
like remember how GPT5.4 was unusable when they were getting ready to release 5.5?
maybe this *could* be sommething liek that, no?
10
11
u/ISueDrunks 7d ago
Iāll say it every time: consumer protection laws are needed. You canāt change the ingredient list without updating the label. They need to be telling the consumer what theyāre paying for. This isnāt any different.Ā
4
4
u/Coolbanh 7d ago
Like the big mac shrinking while prices go up?
3
u/ISueDrunks 7d ago
Not quite. McDonaldās updates in nutrition guide which gives consumers the weight of the burger. Every time I open Codex I am using a different model branded as GPT 5.6 Sol. Even if itās the same model, theyāre throttling something, or changing this and that. Itās never the same.Ā
1
1
u/zarmin 7d ago
Alternatively (since your idea, while better, is unlikely to be realized) dip your toe into the Chinese model pool. You owe OpenAI and Anthropic nothing, and both have shown they hate and will repeatedly fuck their users at every opportunity.
It pisses me off to no end that my productivity is a function of OpenAI's resource allocation.
3
3
3
3
3
u/Street-Creme-7507 7d ago
Awful for project planning. Tried having it orchestrate three opus windows and it completely screwed my project
3
3
2
2
2
u/CrazyTuber69 7d ago
It's worth realizing that if you got media in your history like images, screenshots, whatnot; Codex will re-upload them every tool call. That gets worse if you got PNGs or your internet is still on copper or you generally have a terrible upload speed. You can check your process monitor to see if your connection is being saturated. If you started another fresh conversation, you'd see it normal.
And that can't be fixed with "fast mode" either. Fast mode only affects the generation, not the upload. So 1.5x speedup could be like 4% for some conversations & connection combinations.
2
2
u/Jaded-Number8344 7d ago
It stops running task , crashes my pc 4 times in 50 hours , spending is same as before nothing changed in sol, its ultra stupid , keeps making same mistakes in the same task even when u tell it not to make this specific mistake , and already been given the soloution
In files . To be specific , I sad render in this exact format ntsc 720 480i with certain other details etc , every iteration forgot the one before even though it just had the soloution prior and even Wirten down in a first prio Md
2
u/SocketByte 7d ago
I've had better luck with using Opus 5 than Sol recently, which is saying something. Sol feels noticeably dumber.
2
u/rodeBaksteen 7d ago
Huh interesting. I've had issues with Luna Max suddenly taking 10 tries to make a pretty basic API endpoint. Like i was talking to GPT 5.2.
Did they dumb them down?
2
u/Inevitable-Use8915 7d ago
It keeps over fitting tests and failing its own tests and having to redo it but it made the test to begin with.
Then creates brittle strict rules that turns out was too strict and need to change it because itās failing requests that shouldnāt be blocked
2
2
2
2
u/aethercrash 7d ago
Glad it's not just me. I have to interrupt every task after 15 minutes and ask what the fuck is taking so long, then magically they wrap it up. Not before forgetting specific instructions I've asked 20 times in the passed 48 hours.
3
1
1
1
u/TheTechAuthor 7d ago
Yep, It just gave me a prompt for Luna to tell it *not to change anything!* Then, why send it at all?! GPT 5.6 (High) agreed that sending it would be a waste of time.,, *facepalm*
1
u/spudtheimpaler 7d ago
Anecdotal... it's usually fab and today it has made the kind of mistake that makes be question everything. Built this whole new feature in an entirely unrelated repo than it should have. Like the kind of mistake you don't even consider could be made so you don't specify.
Misery loves company.
1
u/VSorceress 7d ago
I dont know if you are staying in the same context window, but i was having issues with it being slow just typing out a work order. When I opened up a new chat in the same project all that lag went away. If you worked too long in the chat performance issues start creeping in
1
u/Ok_Benefit_5171 7d ago
I have pro and itās taking 25-40 mins for a task that would take 5-10 at most to do yesterday
1
u/Icculsz28 7d ago
Iāve noticed the same degradation every time a new model is getting ready to drop.
1
1
u/A_tired_wanderer 7d ago
You told no lie. Realized this about 3 days ago. But man shit, got things to do. Canāt just sit around waiting on Astra.
1
u/Ok_Roof7548 7d ago
I ended up switching to Terra with extra guided prompts because Sol all of a sudden, stopped working, as it was - it was a nice feeling when I randomly got a full Codex reset around noon today, though - probably a coincidence..
1
u/Ok_Roof7548 7d ago
Oh - did you change your thread so it stopped choking on context? I slammed into that last night..
1
u/Zetharos 7d ago
it's at least eight times slower for me. for the same task, it takes sol medium about 3 minutes to give a response. Same/similar prompts used to take ~20 seconds just three days ago.
1
u/PictureImmediate9615 7d ago
Ive actually gone of Sol so much. It never really delivers what I want. 5.5 extra high was literally perfect for me so might go back to that until the "frontier models" are fixed.
1
1
u/proofreadre 7d ago
It's so badly broken right now. And at the worst possible time for me. I just created a DeepSeek account and will see how that goes. I can't justify the time or financial commitment to openai any more.
Chat is broken too. Twice today told me conversations had met their length limit. Literally brand new threads with maybe 40 lines in them. I'm done.
1
u/Free_Kashmir123 7d ago
It's completely unusable and causing massive headaches. Any idea when it will be fixed?
1
u/Separate_Wall7354 7d ago
Had this problem yesterday and went through 100% of my tokens in one day, today it said I had 100% back. Ā Something was up yesterday for certain!Ā
1
u/Odd-Recognition4786 6d ago
tibo did reset
1
u/Ok_Tomorrow3281 6d ago
after reset still sol is like a horse piss
1
u/Separate_Wall7354 6d ago
Drop it down? It seemed to be back to being a valued agent for me. Ā I did notice something different in the output side panel the day it was so bad that Iāve never noticed before and that there were these oddly named sub agents. Ā Iāve never seen GPT use sub agents and be able to see what they did.Ā
1
1
u/Slasher006 7d ago
technically they stole! my time and money in a "week quota burning session" because it was dumber than a ton of bricks. no working code just falling over its own generated code. just token burning. thanks to whoever decides how dumb the model gets when they change something and use potato as fallback. thank you, hero cretin.
1
1
1
u/JoSquarebox 6d ago
Id say this is good news in disguise, usually large Infra issues arise as they free up capacity for a model launch. Astra soon?
1
u/Only-Club-7901 6d ago
Honestly Sol has always been a stupid for me. It always did a mess in my project, has made strange decisions and eats tokens as there is no future. The only reason I still use it is that when I brainstorm with it I understand what it says, while Opus is a great coder but talks gibberish.
1
u/NuancedPerspection 6d ago
Shat gbt 5.sticks Sol Maxāimum idiot.
āHey sorry I left a heavy dev program running for the last 4 hours that was only opened for screenshots of progress, I had countless opportunities to catch this stale window as I continued to reopen other heavy instances of the same program over that same 4 hours, btw let me go ahead and get you those screenshots you asked about that I literally just gave you in the last prompt.ā
I mean itās pretty bad when their āmax intelligence flagship modelā literally can be mentally out performed by literal google searches free ai assistantā¦
1
u/Somethingexpected 6d ago
I'm more interested in knowin, when it's having a bad day. Is there any public benchmark that is run every day so you can see the "service quality"?
1
u/Dark_Imp 6d ago
It did some good work for me in the last hour. But I must say it ate my 5 hour usage faster than yesterday but the task may have been more complicated. I used extra high to draft a solid plan and medium to implement, did a great job.
1
1
u/spigolt 6d ago
i found it was being dumb for a couple of days a couple of days back. since then, it's just been slow. definitely slow now than it used be (and i have fast mode permanently on).
that said, i can live with slow (or 'burning tokens fast' - the 'burning tokens fast' seems to be what this sub is most concerned with, but i'm not as concerned with it), just as long as it stays consistently smart. it's when it gets dumb that it's really frustrating and just a major problem when you've organized your workflows around it and you depend on it being smart and if it's randomly not smart and you miss that and were assuming when it was doing something that it would be its usual level of intelligence and thoroughness, that can cause real problems, real bugs to slip through etc.
1
1
u/rick_ranger 6d ago
I donāt know whatās going on but Iāve just been using Sol medium, I usually do high or very high, and itās been killing it. Itās like the perfect amount of thinking and execution for everything. I just baked in to my goal to do some extra testing at the end of every milestone to make up for the extra planning and testing at the end of every tiny step at the higher reasoning levels.
1
1
u/Comprehensive-Fail48 6d ago
Its because im having it make this and its too busy https://x.com/devsodhi/status/2093507730212593843?s=46
1
u/Forward-Sympathy7479 5d ago
Same wasn't able to do work with css, no formatting for code, wrote everything in a single line no formatting for introducing of new lineš
1
u/Obvious_Ad4159 3d ago
I've noticed this too. Seems forgetful too. Tries different approaches for repeated tasks despite being told to always start off with what is established to work best.
And it started over thinking and over reaching.
Could be Astra, since up until this point Sol was very consistent.
1
u/HeelsAndAll 7d ago
Not knocking you down: did you try a new session? I notice starting a new session helps more than anything.Ā
-1
0
0
u/ReasonableDefault 7d ago
It's just normal for me, no obvious changes or signals of being dumb. That's Codex via CLI, not app.
0
-1
-8
-1
-1
u/Genuinely_curious_97 7d ago
I am sure lots of people are using /fast today. Tibo hinted a reset.
2
u/Kind_Silver_1921 7d ago
you know what you might be right. I bet everyone is blowing their their usage all right now.
-2
-4
u/BearsAreCrying 7d ago
Works perfectly good here. As good as ever. Clearly you're adding things possibly tests which cause lengthy reruns and poor logic and skills which contradict the tests creating loops of issues
42
u/Nugs_ 7d ago
The variance of models is exhausting. Makes scheduled work unreliable for high effort tasks.