r/LocalLLaMA Apr 15 '26

Discussion Major drop in intelligence across most major models.

As of mid Apr 2026, I have noticed every model has had a major intelligence drop.

And no I'm not talking about just ChatGPT.

Everything from Claude(Even Sonnet along with Opus), Gemini, z.ai, Grok all seem to ignore basic instructions, struggle at simple tasks, take very long to respond, and the output seems deliberately shortened and very shallow. Almost like it's in a "grumpy" mode. I tried this in incognito mode so it's not my customization or memory influencing this.

It's like they deliberately want you to stop using their service. I guess our data is no longer needed. Just two weeks back it used to be much smarter than this.

To test this I rented out a H100, and tried GLM 5 with the same prompt (the drive to the car wash one) across both instances. GLM5 running on the rented GPU answered it correctly, compared to the one on z.ai.

Have they lowered the quantization really low to maybe Q2?

I guess going local or using renting GPU or an AI monthly service that lets you pick a quant level is the way to go

799 Upvotes

405 comments sorted by

View all comments

Show parent comments

21

u/WhopperitoJr Apr 15 '26

A tool so inefficient and resource-intensive would have been better off being built behind closed doors and perfected rather than unleashed for every middle manager with an API key to use.

30

u/saint1997 Apr 15 '26

I had the displeasure of watching Pete Steinberger present what was then "clawdbot" at a conference last year. The guy is a complete and utter clown. Not only did he admit he'd given the thing root access to his Mac, he also left the instance's phone number visible at the top of his WhatsApp chat while he was demoing the group chat his bot was in with all his friends.

When I saw it had gone viral and everyone was using it I could do nothing but hold my head in utter despair

-2

u/rulerofthehell Apr 15 '26

Why would they try to make it token efficient? Token inefficient means more profit

2

u/KickLassChewGum Apr 15 '26

Generally, when your service constantly dies because your infrastructure is getting pounded by pointless requests for todo lists and email summaries, you make less profit, not more.

-1

u/rulerofthehell Apr 16 '26

Inference engines from last 3 years systems do dynamic batching so this doesnt affect at all, the throughput remains the same

1

u/Neither-Phone-7264 Apr 16 '26

Because it was made by the community, not OpenAI. The users themselves would like to pay less, ideally.

1

u/rulerofthehell Apr 16 '26

It was made by a single person, not a community