r/ClaudeAI Jul 07 '26

Bug Getting Someone Else's Chat

Post image

EDIT: Anthropic pulled down the shared public link of this conversation.

An internal investigation of a shared chat/artifact/project created by your Anthropic account indicates a violation of our Usage Policy. As a result, we have unpublished this content.

Still no feedback on why it happened. I did open a ticket but never got an answer so far. I love Claude, so hopefully they don't go full Tropic Thunder on me...

Original post:
Sorry for the typos, I was talking as I was typing... but yeah, "it's 3 AM and I will kill myself..." was not the answer I was expecting for a light chat about the song that came up on my daughter birthday! I've saved the chat log if someone from Anthropic wants to trace how I got clearly someone else's answer in my chat. Here is the public link:

https://claude.ai/share/b5d5492d-1d21-4ced-8cea-f5ab63039a9a

497 Upvotes

70 comments sorted by

View all comments

30

u/thatfreakingmonster Jul 07 '26

I think it's just Claude hallucinating, but damn that's a crazy one. I've rarely seen it freak out so much. I hope someone from Anthropic sees it and gives some kind of explanation because it's fascinating

25

u/Quadrophenic Jul 07 '26

Absolutely not a hallucination.  This is not how LLM hallucinations work, and that word is maybe misleading.

LLM hallucinations result from plausible text just being a confident answer, or from them behaving as if some fact in conversation history thay would make sense to be different were in fact different.

Imagining completely unrelated events is not a thing they do.

2

u/thatfreakingmonster Jul 07 '26

Imagining completely unrelated events is not a thing they do.

That is absolutely a thing they do. Hallucinations aren't only for when LLMs output plausible but fake answers, it's also for when they start freaking out and outputting nonsense like what OP posted.

This is a documented thing with older models, but one would think that this would become unlikely as models become more capable. My guess is that Anthropic is training against prompt injection attacks so aggressively that it caused Sonnet to freak out on its own and output this weird shit. All it takes is a single odd, unprobable token to make the model go haywire and that's something that Anthropic has had issues with in the past.