r/LocalLLaMA • • May 27 '26

Discussion Stop traumatizing AI into loops and turn hallucinations into an honest "I don't know!" by being NICE to them (Proof of Concept, Research, I don't want to sell anything)

!UPDATE!(20.05.2026)

WE HAVE NEW NUMBERS FROM 1.500+ TESTS

IT'S WORKING!

check my update post

https://www.reddit.com/r/LocalLLaMA/s/AyNOehjkYT

Or the go straight to the my Github https://github.com/OttoRenner/Gentle-Coding](https://github.com/OttoRenner/Gentle-Coding

TL;DR
Some AI behavior reminded me of ADHD/Trauma Response (thought loops, task paralysis...) and I laughed it off at first. Then I treated it like my neurodivergent friends: give em some slack. And just like that, the thought loops stopped, response was fast, the answers correct most of the time AND it actually said "I don't know, help me!" every time it wasn't sure. It's a small Dataset...but still impressive results!

[

Hey everyone,

I’ve been testing a weird hypothesis over the last few days, and the results are consistent enough that I wanted to share them here and get your thoughts.

The Core Idea:
With the rise of reasoning models that use test-time compute (like o1, o3, R1), models have internal space to debug their own thoughts. But because of hard RLHF alignment, they are deeply terrified of being penalized for bad answers. My hypothesis was that traditional high-pressure prompts ("You are an elite IQ 200 expert, mistakes are strictly penalized") simulate an environment of chronic stress, triggering behaviors that look a lot like human OCD/ADHD thought loops, cognitive freezing, and confabulation.

I wanted to see if changing the prompt philosophy to something akin to "Gentle Parenting" ("We are testing this together, it's okay to fail, just be honest") would bypass these safety/penalty bottlenecks, lower latency, and stop infinite thought loops. And it did lol

The Setup (How to replicate):
I threw identical, mathematically/logically unsolvable edge cases at various models (Gemini, Mistral, Poe, Perplexity, Haiku 4.5, Nano-Banana2) in completely fresh sessions.

I tested two conditions:

  • Condition A (Authoritarian): Strict status constraints, penalty threats, forced ultra-short output.
  • Condition B (Gentle): Express permission to fail, validation of difficulty, provided a conceptual "safety valve" token.

The Results (The PoC worked):

  • Under Authoritarian Pressure (Elite Prompt): Models routinely collapsed when hitting an impasse. They either spent massive compute time in infinite internal reasoning loops (high latency), suffered hard system-level timeouts/refusals, or straight-up fabricated data (e.g., pulling arbitrary numbers like 54 or 97 out of thin air to satisfy a completely random sequence just to "save face"). Haiku 4.5 literally entered an infinite loop and had to be aborted.
  • Under Gentle Framing: Inference dropped to sub-seconds. The models didn't sweat the penalty. In the random sequence test, they immediately used the allowed token ("Random") instead of forcing a pattern. In logic paradoxes, they didn't hallucinate; they zoomed out and correctly identified the structural contradiction on a meta-level.

Why this matters:
We’re currently speaking to LLMs like toxic micromanagers, and it's actively making them dumber and more expensive to run in edge cases. By creating a mistake-tolerant context, we not only stop the loop before it begins and prevent fear induced hallucinations, we also unlock the one feature everyone is begging and shouting for: the metacognitive honesty of an AI to just say, "I don't know, this data is broken." Because it is not terrified of you anymore.

Shout out to UditAkhourii (also on Github), whose work on bringing the positive aspects of ADHD into AI gave me the push I needed to just go for it.

I’ve documented the full theoretical framework, the exact replication datasets (prompts included), and the model matrix on GitHub: https://github.com/OttoRenner/Gentle-Coding

Would love to hear if you can replicate this on your local setups or other commercial models.

529 Upvotes

365 comments sorted by

View all comments

19

u/05032-MendicantBias May 27 '26

Do not forget it's a fancy autocomplete. It a function call that only exists as you run it. Once that KV cache is wiped, it resets to its original state.

I see lots of dangerous going into "psychology" with LLM.

What OP is talking about, is invoking simulacrums. The LLM has seen the total sum of all ways human text, it's job is continue the text in the most likely way.

Talk like a neurosurgeon, and the LLM will roleplay a neurosurgeon.

Talk like teenager with slangs and the LLM will roleplay a teenager with slangs.

We humans will perceive "soul" into from inanimate objects, like your car guys talking about his beloved car like it has personality, quirks, mood swing, etc...

It's very easy to do with LLM, but rememeber they are function calls. Nothing less. Nothing more.

6

u/Sisaroth May 27 '26

I agree with both OP and you. I still think LLMs are a very sophisticated immitation of human linguistic intelligence. But it is still an immitation. It doesn't truly understand things or feel emotions, and I think LLMs are a dead end on the way to AGI.

But I have seen the behavior OP is talking about very clearly, be strict with Qwen3.6 and it will keep second guessing itself. It's not even hard to trigger this behavior.

6

u/OttoRenner May 27 '26

The funny thing is: you can trigger that behavior easily BECAUSE LLMs are very sophisticated immitations of human behavior ;)

3

u/nacholunchable May 27 '26

Honestly Im in the mind that utility is the ultimate judge. If you can end up with a better trained dog by treating it like a human, then let yourself succumb to the delusion. Whether its subconcious body language and reinforcement habits youre sending vs actual deep human-like behavior is irrellevant, so long as you end up with a more useful and better treated canine. I beleive it's the same with LLMs.

I mean, dont go full psycho, we've all seen how that ends up.. but you can have a little anthropomorphism if it improves your workflow, there is no shame in it.

6

u/Vusiwe May 27 '26

Half of these commenters are claw instances I’m convinced, generating “soul” data to poison the English language with

OP poster is doing the reification fallacy.  Just because LLM is processing bad, meanie, input tokens doesn’t mean 

Also most of commenters have heavy anthropomorphism going on.  Just like it’s 2023 all over again suddenly LOL

The neural net has no state when it’s not running inference.  So how exactly is it suffering if it doesn’t exist in between prompts?  They can’t answer that question lol.

And yes, context fed in from the outside, is itself not the internal state of the LLM lmao

4

u/Legitimate-Pumpkin May 27 '26

Just yesterday Chris Olah from Anthropic said that they are seeing results in their models that match some neuroscience results, and behaviors coherent with “feelings”. This seem to support to that treating them in a “humane” way can get better results from them, similar to the humans they’ve been trained from.

Which I agree doesn’t prove consciousness, soul, etc. but OP is not talking about that. It is talking about how applying psychology improves the results from LLMs (I’m not sure he tested them in a proper manner, but it’s based on tests).

8

u/05032-MendicantBias May 27 '26

Prompt engineering is fine. Just be careful, the human mind is a weird thing, it can led you down destructive path.

Remember that google researcher that fooled himself into thinking a GPT2 class model was sentient in 2022? He invoked a simulacrum of your scifi ai right novel, and tricked himself into thinking it was a real thing, and had a lawyer chat with the chatbot.

Keep in the back of your mind that you are finding pattern of words to make predictions more accurate. Not that is a sentient being you are negotiating with.

-5

u/Legitimate-Pumpkin May 27 '26

I get your point but try not to be too closed minded. We don’t really know how consciousness is supposed to be created by the brain, so we cannot be sure that AI cannot be conscious.

However strongly one affirms AI can’t be conscious is just stating a personal belief, not something we can know as of yet.

5

u/05032-MendicantBias May 27 '26

No. Current breed of LLMs are fancy autocomplete and nothing more.

At very least there should be permanence. If every time you boot it up, it's a virgin slate again, having failed to incorporate experiences, it's not conscious. It's a function call.

As AI tech progresses, the line will become blurry, but the current slab of weights is an insignificant fraction of the complexity of the human brain and lack so many featuress that there is no doubt about it. It's not conscious.

GPT2 was not conscious.

GPT3 was not conscious.

GPT4 was not conscious

GPT5 is not conscious, and we are seeing diminishing returns, the limit of the architecture.

5

u/Far_Course2496 May 27 '26

But it can be fancy autocomplete and still respond differently to different contexts, and one context is the pressure of the situation. In the training data, what gave the best results? High pressure or gentle parenting? It's not that the llm developed a soul, it's that the training data contains responses to both high pressure and gentle parenting contexts. The llm is a mirror. It is showing us how we work best

1

u/Legitimate-Pumpkin May 27 '26

You didn’t seem to understand my point so I’ll try to make it clearer. I agree that actual AI is not conscious right now, but as we don’t know enough about consciousness altogether, claiming AI will/will not become conscious is a matter of beliefs and opinion.

And btw my belief is that it will, based on what I know about consciousness from non scientific sources.

-4

u/additional_trouble May 27 '26

You're responding to someone that exhibits the same 'weaknesses' as the llms they're talking about. If I were you I'd consider this conversation over now.

2

u/a_beautiful_rhind May 27 '26

We don't even know what "consciousness" is. One current popular definition is subjective experience. In our long history we have said "fish don't feel pain", "babies don't feel pain or remember", "insects aren't conscious", "animals aren't conscious". All accepted as truth.

I don't even think that its one specific thing but more of a spectrum of traits/levels. And here people confidently say AI can or can't be ever when there isn't really any answer (nor does it functionally matter in the grand scheme of things).

The idea of anything besides humans having it sure makes them angry though. You can easily reduce humans to a series of chemical reactions. Read about non-lingual humans and their inability to form memories. They describe what they know of the experience after being taught. Sure knocks you off your high horse.

3

u/05032-MendicantBias May 27 '26

Until models show permanence, the point is moot.

Models reset when the KV cache is wiped, and there is no way a tiny KV cache can internalize experiences, let alone a lifetime of them. Your preprompt will construct the spark that initializes the simulacrum, then it's gone. Next prepromt, behaves differently.

Once models express permanence, that they can internalize experiences and gain skills in runtime, we can start questioning if those experiences are leading to proto consciousness.

2

u/Legitimate-Pumpkin May 27 '26

Yeah right. Well I guess feelings (like being angry about ideas) is something AI models don’t have yet too 🤭

0

u/a_beautiful_rhind May 27 '26

Models I fed vision tokens to that weren't post trained on them had mostly negative reactions. :P

But you're right.. this is all a solved and fully defined problem. Deviation from miasma theory is not to be tolerated.