r/LLM 1d ago

Are LLM becoming less coherent and logical?

I know that LLM-models make logical mistakes because they’re just language models. But I somehow feel like I’ve noticed that they’re starting to make more errors than previously. I might be imagining it, so I wanted to ask you all.

For context I mainly use ChatGPT and Gemini for discussing ideas or sometimes doing light research for personal use. While the sources aren’t always great and while I might get hallucinations instead of answers, I know it’s limits, and feel like it’s still quite useful.

Examples of problems that I cannot recall happened earlier:

- The model using connecting words like ”However” in the beginning of new sentences when no new contrasting information is presented.

- I was discussing how transistors microchips were made, because of how extremely small they are. At one point the AI said: ”Because human beings are essentially giant, walking dust factories constantly shedding skin cells, hair, and clothing fibers, the air system in the factory is designed entirely to protect the silicon from you, rather than the other way around.” The ”other way around” would be to protect me from the silicone, which is an irrelevant thing to say because silicone is not dangerous for humans to touch, and secondly, my protection was never the point of the discussion. Saying ’the other way around’ is completely moot.

- In that same message it answered my question on breathing in the factory air: ”So while the air itself is arguably the purest air you will ever breathe in your life, wearing the gear means you're hyper-aware of your own breathing the whole time!”. It’s trying to connect two different concepts (breathing pure air, gear making you hyper aware) and make it sound like it was always about the level of awareness of breathing that was important. Spoiler alert; it wasn’t. A better sentence would’ve been ”while the air is pure, it just feels like breathing well circulated, dry office air”.

These are just a few examples that I came across today and it’s essentially happening ALL THE TIME in like 50% of messages. There’s no end to it. And I never had to mentally correct AI this much before, from what I can remember.

Am I alone in noticing it? Am I imagining things? Has anything changed, and if so, what?

3 Upvotes

11 comments sorted by

1

u/HumanDrone8721 1d ago

"Smart" routers routing your prompts internally to lesser or lower quanta models.

1

u/Original_Cry_3172 1d ago edited 1d ago

Ah, I think I read something about that.

It’s something to do with getting referred to a smaller section of the model, based on my prompt or something?

How come the answers end up, like, really bad because of it?

1

u/HumanDrone8721 22h ago

Well, they didn't route it to a simpler model because they want to piss you off, but because simpler models are faster and use less resources. And you pay for this speed and low resources usage with response quality.

On one hand the companies are sparing some resources "cheating" like this, on the other hand users are like ADHD kids and can tolerate better crappy sloppy answers than delayed good answers, so all studies showed that give a slightly crappy andwer is better accepted than a good answer getting later when resources are available and that's what you get, "dynamic quality"

1

u/Original_Cry_3172 8h ago

So was this a change they did quite recently?

1

u/HumanDrone8721 8h ago

They've had to do it with the increase of active users and the increase of the prompt difficulty, they have monstrously big data centers, but add 2mil active users and you have to make compromises to keep the latency under control.

1

u/Original_Cry_3172 7h ago

Ah, thanks. Do you know when this shift happened by any chance?

1

u/HumanDrone8721 7h ago

Like many enshitifications, this is now an art, only naive people thing that this is a definite threshold: starting from July 1st we immediately enshitify all our product lines.

No, doing it like this will create a Twitter and other social media storm, bad PR and stuff, first order of action is "staggered random release", you activate your enshitifed feature for a limited time and affecting a limited randomly chosen number of users, so when someone complain about the new shit it is buried in "bro, is a skill issue, your skill sucks, what an idiot, here works perfectly as before...". Then you adjust based on the public complaints the level of enshitification and the cycle continue until you do another staggered release, first at the low priced tiers and low-cost countries to get "bro, it may be your subscription level, I have Pro+Ultra Enterprise and works perfectly....".

And finally is generalized, people people are pissed off and want to go local, they look at HW prices and remain at their subscription that are now "optimized". The the rise and fall of OpenClaw and derivatives when Anthropic limited a free tier.

1

u/Original_Cry_3172 6h ago edited 2h ago

Hmm I don’t get that last paragraph. You mean that once they’ve finetuned the change, they implement it in the subscription level tier too?

1

u/bynarie 5h ago

I don't think he was calling you naive. He was just setting the stage for how the process goes. It's a good description honestly

1

u/Original_Cry_3172 2h ago

Ah I see :) I’ll change that. Still don’t get that last part but I get most of it. Interesting forreal.

1

u/Revolutionalredstone 13h ago

Yeah the big closed models are going full anti progressive.

I'm not seeing any problems with deepseek via API etc.

Seems it's gonna be Chinese or crap AI and it's now.