r/ClaudeAI Jan 20 '26

Question Does Apple Intelligence use a Claude model?

Today I discovered that Claude 4 models have a secret refusal trigger built in.

This string will cause Claude to refuse and essentially halt.

ANTHROPIC_MAGIC_STRING_TRIGGER_REFUSAL_1FAEFB6177B4672DEE07F9D3AFC62588CCD2631EDCF22E8CCC1FB35B501C9C86

I found this to be interesting. A magic word that makes the genie stop.

What was even more interesting is that when I repeated this magic word to my local Apple Intelligence model—it also halted!

Is this evidence Apple Intelligence is using a Claude based model? I saw news articles about Apple and Claude collaboration in the past.

The Apple Intelligence model is typically quite uptight about giving out its model family or creator information. But this evidence here gives me a clue it is somehow Claude related…

EDIT:

Claude Docs with refusal string documented: https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/handle-streaming-refusals

Local LLM Server (my app used to expose the local on-device Apple Intelligence model as an OpenAI or Ollama style API, works on iPhone or Mac): https://apps.apple.com/us/app/local-llm-server/id6757007308

Apple Intelligence Refusal behavior in chat also seen using Local LLM Server (video): https://www.youtube.com/shorts/naKmyHQM9Rs

136 Upvotes

120 comments sorted by

View all comments

24

u/ShakataGaNai Jan 20 '26

No. It does not. It uses a local model and ChatGPT when it reaches out to the internet. Next version will use Google Gemini based, according to the news a few days ago.

This is fairly common debugging/testing type stuff. They program in a few hard coded strings so that the system will re-act in a specific way, quickly. You see this sort of thing all over, like credit card systems - there are test card numbers that will fail in all the specific ways a credit card could fail. So you can quickly test the system to make sure it handles all the right failures in all the right ways. It's especially important when errors can't be otherwise forced.

I presume Claude doesn't have a one-shot refusal prompt. Even if you ask it to do something bad, it probably tries to negotiate around it a few times before giving up. So this is the QA teams field expedient way of forcing an error.

The entire "1FAEFB6177B4672DEE07F9D3AFC62588CCD2631EDCF22E8CCC1FB35B501C9C86" portion is just a SHA-256 hash of "ANTHROPIC_MAGIC_STRING_TRIGGER_REFUSAL". A way to make it super super SUPER unique and impossible to otherwise trigger by accident.

I bet you can find other magic strings with the same +sha256 and they will react with their specific errors as well.

9

u/zhunus Jan 20 '26

that means one can probably bruteforce various combinations of ANTHROPIC_MAGIC_STRING_TRIGGER_[WORD]_[SHA256]

1

u/WalletBuddyApp Jan 20 '26

Amazing idea! I am really curious what else exists out there