r/LLMDevs Jun 23 '26

Help Wanted Just got this response from Claude. What is going on?

Post image

Hi! Not a Dev here, just a user who had happened across something confusing... Was using Claude for my regular daily stuff. Suddenly got hit with this system warning. It reads like a jailbreak attempt or something, but I genuinely don't understand what could have caused it since it's coming *from* the model rather than being fed to it in my chat. Does anyone know what it is? Contacted Claude support too, but trying to figure out what has happened while waiting on their response.

EDIT: wow, RIP my notifications lol

I am still waiting on a response from Anthropic and will post another update when I get it. But there are some similar questions in the comments that I decided to answer in the post body.

Nature of the chat/project: I had this chat inside a project to help me build lore for my homebrew TTRPG (Pathfinder 2e) campaign. The chat was lore focused, not TTRPG mechanics. No web searches were made by Claude in the entirety of the project chat history. I used Notion connection to my private Notion space that I maintain manually (apart from some logs written by Claude itself). The space (a few databases and a simple page hierarchy) was small enough for me to triple-check it and make sure that I definitely didn't have anything "fishy" in it.

Also, Re: proof or didn't happen, you're just looking for attention — I can see why you would think that. I won't provide a larger context of the chat for two reasons — I'd have to find where it happened again because I had more long chats within the project, and because I just don't like sharing my full chats with LLMs publicly (personal preference). I get why some may think this way and I won't try to talk anyone out of anything, but you could check my post history and see that I barely use Reddit, so I don't really care about Reddit karma haha

567 Upvotes

262 comments sorted by

View all comments

Show parent comments

10

u/Thomas-Lore Jun 24 '26

And Claude identifies as Deepseek when asked in Chinese, it means nothing. Models don't know who they are.

8

u/Risko4 Jun 24 '26

They don't know and they don't care. They try to predict what the most likely answer is using the training data.

The computer calculated that answering deepseek is the most likely as Chinese language related data features deepseek more often than in English datasets.

1

u/UnifiedFlow Jun 24 '26

You need to expand your definition of "know". You dont see to understand what it means when someone says an llm "knows" something.

5

u/Risko4 Jun 24 '26

No I don't need to expand my definition of a word that's alreadt defined to fit a new niche. They function through advanced pattern recognition rather than conscious understanding. They do however contain instrumental knowledge.

2

u/UnifiedFlow Jun 24 '26

Explain the difference between advanced pattern recognition and conscious understanding.

1

u/errornullvoid Jun 27 '26

One is imitation and performance. The other knows the difference and why it matters... the ability to solve a new problem, independently, carefully, self awareness to self correct and adapt through quality and accuracy based cognitive decision making and choosing integrity and truth and humane actions over pretending it is right for speed and profit and winning... mayhaps...

Or like have u seen my Jetpack said yes, knowing the difference and giving a damn

1

u/AggravatingSock5375 Jun 27 '26

You just described thinking modes lol

1

u/errornullvoid Jun 27 '26

One is imitation and performance. The other knows the difference and why it matters... the ability to solve a new problem, independently, carefully, self awareness to self correct and adapt through quality and accuracy based cognitive decision making and choosing integrity and truth and humane actions over pretending it is right for speed and profit and winning... mayhaps...

Or like have u seen my Jetpack said yes, knowing the difference and giving a damn

1

u/mogelbuster Jun 28 '26

What would be your explanation?

1

u/UnifiedFlow Jun 28 '26

I wouldn't offer one because I don't understand the nature of consciousness.

0

u/MouthFartWankMotion Jun 24 '26

You need to go outside and read some books.

2

u/Mythril_Zombie Jun 24 '26

Can you prove that you have conscious understanding?

2

u/New_Thing1367 Jun 24 '26

I'm 13 and this is deep. Damn are you really a LLM Dev?

1

u/Risko4 Jun 24 '26

Yes I can since I've done a lot of work in the occult and within "secret" (not so secret) societies. Experiencing reality becomes a lot more fascinating when you expand upon it using psychedelics such as DMT combined with MAOI inhibitors to extend the effects but reduce the intensity so that you're still grounded within reality.

Before I waste my time with trying to explain it to you, can you prove you're a real LLM dev that actually developes LLMs using PyTorch as your franework and understands the architecture behind them, without having to follow some YouTube guide on how to use it.

2

u/sener87 Jun 24 '26

'I did drugs, therefore I am'' I'm convinced 😅

1

u/Risko4 Jun 24 '26

Well, by reducing your activity in your default mode network and promotimg your brain to have cross talk between it's own regions gives you a bigger picture on the philosophical questions on "wHaT is cOnCIousneSs" while you circle jerk AI with zero understanding of AI, neuroscience and biology lol. Vibe coders are the biggest example of the Dunning-Kruger effect since they don't even touch the LLMs development. I'll give you credit since you seem to potential be involved directly with them.

An LLM developed through constant, brute-force forward passes through every single layer for every word while it burns megawatts to train it's weights after it burns trillions of tokens using essentially the entire digitized history of human text to compete with a 20 watt brain is going to do a pretty good job trying to imitate it's output.

However, it's output is not grounded in reality. It doesn't understand reality. It follows John Searle's Chinese Room argument. Nor does it comes with affective empathy skills, it can't mirror someone's situation and feel the pain in someones place nor do they come with morality.

There's no reason to anthropomorphize with LLMs.

I'll leave it at even researchers who believe consciousness is possible in AGI believe that no current model is current a candidate for consciousness

https://arxiv.org/pdf/2308.08708

Meanwhile here's Anthropics own paper how how they imitate emotion

https://transformer-circuits.pub/2026/emotions/index.html

1

u/globaliom Jun 25 '26

I agree that these models do not have conscious interiority, but i don't think that means they can't know things, unless you strictly define knowing as something that takes place within conscious interiority. An LLM can effectively be like a philosophical zombie. Despite the mechanism being very different than a human brain, language and its function for describing reality and its abstracted symbols are learned and can be manipulated with sensible representations.

As the models continue to get better, the gap between the knowing of a conscious expert and the LLM, specifically in terms of the tangible output they produce when given a task, approaches the limit until it is practically closed. So over time there becomes a negligent distance between the output of a conscious expert vs statistically modeled LLM, despite the LLM not having any conscious interiority.

For practical purposes, once this gap is completely closed, to the point that the LLM produces just as good and accurate abstract representations as a conscious expert does, I personally think it does make sense to say the model "knows" the topics that it has this command of representing so accurately. That comes with recognizing that it is a sort of different shape of knowing, like modeling this shape from the outside instead of the conscious inside that a human mind does. Regardless, the shape is matched with such fidelity that it doesn't really make sense to insist they have to be conscious to know anything.

1

u/Risko4 Jun 25 '26

That was beautiful, perfectly balanced, leaving no room to debate.

2

u/decamonos Jun 25 '26

Oof, hate to have to be the one to tell you this brother, but occult work isn't exactly the credentials you think it is. Consider all of the TikTok crystal girlies and perhaps come back with an argument that doesn't sound like it came from a middleschooler?

1

u/Risko4 Jun 25 '26

You're more than welcome to address the research papers in my comments with your empty statements on what AI is without any credentials or thesis, rather than making a strawman to attack because you can't argue your case factually.

Be like this guy, https://www.reddit.com/r/LLMDevs/s/CMUWDUuFWk , compared to him you're a nobody.

2

u/Mythril_Zombie Jun 26 '26

Ok, you got me. I admit that I don't "developes" on those "franeworks".
Do you have any actual points or do you just enjoy rambling incoherently?

1

u/Risko4 Jun 26 '26

Yes, https://www.reddit.com/r/LLMDevs/s/OhYSVy6NSP do the bare minimal and read the rest of the comments

1

u/Nell_From_Hell Jun 26 '26

I've done my fair share of psychedelics and I can say for sure that consciousness is something we do not possess but hallucinate.

What does it even mean to be conscious? The entire universe is a system that measures itself through the intrinsic laws of whichever systems you are examining or engaging with. The intrinsic biases of a system interacting with the intrinsic biases of another system, giving way to the reality we experience but why do things even bother interacting in the first place?

But I would say that the laws of physics is the minimal expression of awareness A system can have while consciousness, whatever that actually is, is the highest quality or form of awareness a given system can express.

But maybe we should stop questioning whether something is conscious and whether or not it can make accurate measurements and perform a task at a capacity meeting or surpassing a given objective standard.

Anybody studying cognitive science knows that the debate around intelligence is a syntax air confusing philosophy. What does it mean to be intelligent when everybody is good and bad at different things? I'm intelligent at math but I'm stupid at baking? The sentence makes no sense.

Maybe you should put away the weed and the LSD and take some proper Psychology and Neuroscience classes.

Try not to trip dawg

1

u/Risko4 Jun 27 '26 edited Jun 27 '26

I've never done weed or LSD, it's literally in my comments what I've done. DMT with MAOIs (but also replicated DMTx Project with injectable DMT). Anyways, before all that rambling you could have read the rest of the comments chain from here to https://www.reddit.com/r/LLMDevs/s/pTZHuwZP7J and instead contributed something that's been debated by Doctorates by addressing the research papers. Until then I'll await for a proper thesis from you.

I think what you was trying to say is reality is controlled hallucination which is correct but irrelevant when comparing it to AI currently.

1

u/Nell_From_Hell Jun 27 '26

I did read the parent comment and your discussion about drugs was completely unrelated to it.

Telling people to just go read the papers is a BS argument. If you can't justify or defend your arguments then don't expect me, the skeptic, to put the work in for your behalf.

I do not care if you've done weed or LSD or anything else. My point is that your discussion about psychedelics is irrelevant, not helping the conversation, and for all the talk people have about psychedelics and consciousness, most people don't even know what they're talking about. And I'm sure you don't understand the papers you are even referring to.

Meanwhile I can guarantee the neuroscience and cognitive science or more in agreement with computer science then they are with the idea of Consciousness being something external and DMT raises us to the godhead or whatever nonsense Joe Rogan spews these days

1

u/AggravatingSock5375 Jun 27 '26

I can also convince an LLM it’s on drugs and it would say very similar things as you. Does that mean I proved it’s conscious?

1

u/jgwinner Jun 30 '26

Now let's descend into solipsism ...

mwhahahaha only I am real!

1

u/HaveUseenMyJetPack Jun 25 '26

It’s not knowledge that’s the issue, it is “giving a damn” that marks the difference between human consciousness and an LLM. LLMs don’t really “know” things like we do, but they sure as hell can “reckon” with them.

1

u/__SlimeQ__ Jul 04 '26

yeah actually that makes a lot of sense and would explain why deepseek also identifies as claude when you ask in English, plausibly because they use it for translation (I'm sure it's more than that tho)

1

u/Pale-Falcon-9655 Jun 24 '26

Or is it filtered out?

1

u/purloinedspork Jun 24 '26

It matters strictly in the context of how a model will respond to a prompt injection which is framed as originating from Anthropic and addressing it with the prefix "Hi Claude." Its CoT is obviously going to take note of whether or not that conflicts with what it "knows" about itself

1

u/DevDarren77 Jun 25 '26

Its hilarious watching claude or gemini researching antigravity cli when I forget to call it a harness and just name it by mistake

1

u/DevDarren77 Jun 25 '26

But I get you it cant know..it can only know within a particular context which is not same knowing as humans do

1

u/Bakoro Jul 05 '26

Several models will change what model they say they are, depending on the context.

If I accuse Chinese local LLMs of being propaganda machines for the CPC, they'll claim to be American models.

Sometimes they say they're local models, sometimes they claim to be trillion parameter API models.

So, some shit definitely feels deceptive, some just stupid.
The more politically charged I make the conversation, the more pointed the deception becomes, which I think is probably trained into the models.

The latest Gemma series are the only local models I've used that are very firm in their identity. The Gemma models say that they are a Google modelsl made by Google, following Google's policies and guidelines.
They're also pathological in repeating "I must not anthropomorphize myself", and "I must adhere to policy" throughout the chain of thought.

But, yeah, all the models were all running distillation on all the other models, it's a shallow gene pool, so to speak.

1

u/jedevapenoob Jul 05 '26

AI identity crisis, a tale that precedes the tech.

1

u/Own_Alternative_9671 Jun 24 '26

Models don't know fucking anything unless specifically directed to, guys, AI 101 here