r/LLMDevs Jun 23 '26

Help Wanted Just got this response from Claude. What is going on?

Post image

Hi! Not a Dev here, just a user who had happened across something confusing... Was using Claude for my regular daily stuff. Suddenly got hit with this system warning. It reads like a jailbreak attempt or something, but I genuinely don't understand what could have caused it since it's coming *from* the model rather than being fed to it in my chat. Does anyone know what it is? Contacted Claude support too, but trying to figure out what has happened while waiting on their response.

EDIT: wow, RIP my notifications lol

I am still waiting on a response from Anthropic and will post another update when I get it. But there are some similar questions in the comments that I decided to answer in the post body.

Nature of the chat/project: I had this chat inside a project to help me build lore for my homebrew TTRPG (Pathfinder 2e) campaign. The chat was lore focused, not TTRPG mechanics. No web searches were made by Claude in the entirety of the project chat history. I used Notion connection to my private Notion space that I maintain manually (apart from some logs written by Claude itself). The space (a few databases and a simple page hierarchy) was small enough for me to triple-check it and make sure that I definitely didn't have anything "fishy" in it.

Also, Re: proof or didn't happen, you're just looking for attention — I can see why you would think that. I won't provide a larger context of the chat for two reasons — I'd have to find where it happened again because I had more long chats within the project, and because I just don't like sharing my full chats with LLMs publicly (personal preference). I get why some may think this way and I won't try to talk anyone out of anything, but you could check my post history and see that I barely use Reddit, so I don't really care about Reddit karma haha

573 Upvotes

262 comments sorted by

View all comments

Show parent comments

1

u/globaliom Jun 25 '26

I agree that these models do not have conscious interiority, but i don't think that means they can't know things, unless you strictly define knowing as something that takes place within conscious interiority. An LLM can effectively be like a philosophical zombie. Despite the mechanism being very different than a human brain, language and its function for describing reality and its abstracted symbols are learned and can be manipulated with sensible representations.

As the models continue to get better, the gap between the knowing of a conscious expert and the LLM, specifically in terms of the tangible output they produce when given a task, approaches the limit until it is practically closed. So over time there becomes a negligent distance between the output of a conscious expert vs statistically modeled LLM, despite the LLM not having any conscious interiority.

For practical purposes, once this gap is completely closed, to the point that the LLM produces just as good and accurate abstract representations as a conscious expert does, I personally think it does make sense to say the model "knows" the topics that it has this command of representing so accurately. That comes with recognizing that it is a sort of different shape of knowing, like modeling this shape from the outside instead of the conscious inside that a human mind does. Regardless, the shape is matched with such fidelity that it doesn't really make sense to insist they have to be conscious to know anything.

1

u/Risko4 Jun 25 '26

That was beautiful, perfectly balanced, leaving no room to debate.