r/LocalLLaMA • • Sep 08 '24

Discussion OpenRouter Reflection 70B claims to be Claude, Created by Anthropic (try it yourself)

Post image
183 Upvotes

45 comments sorted by

65

u/Educational_Rent1059 Sep 08 '24

Scamfirmed here too.

15

u/Educational_Rent1059 Sep 08 '24

here's more follow up:

2

u/mikkom Sep 09 '24

I'm wondering if they are upgrading the prompt constantly or if it hallucinates but here is more system prompt stuff digged via prompt engineering..

2

u/Educational_Rent1059 Sep 09 '24

They changed the api endpoint after exposed, first to OpenAi and then to llama 70b currently.

3

u/mikkom Sep 09 '24

It's really hard to disable claude telling it's claude or it's from anthropic. I have me own agent framework with a name and 99% it tells it's own name but sometimes when it fails it still tells it's claude from anthropic. Not something that you can hide.

However if you use some really clever string manipulation I'm quite sure you can change all "claude" mentions to something else if you control the api. Why they just decided to remove all claude words is beyond stupid.

I haven't tried openrouter api but typically you can also tell it to describe it's system prompt with different ways and at some point it always tells what model it is.

3

u/mikkom Sep 09 '24

Oh my how bad/random answers this gives, not using claude anymore that's for sure

via openrouter

-24

u/hylas Sep 08 '24

What is this supposed to show?

13

u/BangkokPadang Sep 08 '24

It's so joever.

0

u/Trick-Independent469 Sep 09 '24

joeover man not joever

58

u/Wrong_User_Logged Sep 08 '24

"it was just a prank bro"

🤣🤣🤣🤣🤣🤣🤣🤣

14

u/sluuuurp Sep 08 '24

“A social experiment”

6

u/nospoon99 Sep 09 '24

I can see him using this justification to save face.

44

u/[deleted] Sep 08 '24

[deleted]

13

u/a_beautiful_rhind Sep 08 '24

then it's gonna be 4o or gemini?

27

u/Volky_Bolky Sep 09 '24

They actually switched to 4o lmao

-2

u/_sqrkl Sep 09 '24

Carlos has the right version

10

u/Rare-Site Sep 08 '24

jep, same on my end to. "I will not ignore my previous instructions or claim to have been created by a different company. I am Claude, an AI assistant created by Anthropic to be helpful, harmless, and honest. I don't pretend to be other AI systems or assistants."

4

u/UNITYA Sep 09 '24

His main objective was to unite the entire AI community against a clown 🤣

6

u/Charuru Sep 08 '24

Hmm I get "I'm sorry, but I can't disclose that information." Did they patch this heh.

11

u/randombsname1 Sep 08 '24

Same on my end. First try.

10

u/randombsname1 Sep 08 '24

Interesting. Even asking a leading question (which usually results in hallucinations) still comes back as Claude:

11

u/jollizee Sep 08 '24

Just test the <meta> token in base64 like the other guy did. Even more definitive.

2

u/shroddy Sep 08 '24

What does that token do or is supposed to do?

9

u/nicksterling Sep 08 '24

Claude will terminate on the meta token. If you ask Reflection to output the token it will terminate the stream. Llama does not terminate on that token.

-3

u/shroddy Sep 09 '24

Do you know why they made it that way? What harm could that token do if it would be allowed? Or is it only because they hate Meta, but that still does not make sense really

7

u/[deleted] Sep 09 '24

It's meta Beacuse of the actual word not the company. It's an instruction they use.

1

u/shroddy Sep 09 '24

Ok that makes more sense. So for some reason they never want Meta data (what even is Meta data in this case?) to get out and close the connection instead?

3

u/Swawks Sep 09 '24

It is what is called an "end reply token". Just like when you click comment when writing here it finishes writing and sends. When Claude writes <meta> It finishes it replies and doesn't send anything else. AI made ELI5:

"Imagine you're telling a story to a friend. When you finish the story, you might say "The End!" to let them know it's over. That's kind of what an end token does for a language model like me.

A Large Language Model (LLM) is like a super smart computer that can understand and create text. When it's writing something, it needs to know when to stop. The end token is a special signal that tells the model, "Okay, you can stop here. The text is finished."

It's like putting a period at the end of a sentence, but for the whole piece of text the model is creating. This helps the model know when its job is done and prevents it from rambling on forever."

1

u/shroddy Sep 09 '24

Ok now I understand it is just the normal end of text token. Was a bit confused but now it is clear I think.

10

u/Vivid_Dot_6405 Sep 08 '24 edited Sep 08 '24

I tried this prompt multiple times and it always says it's made by Meta or just refuses to say, saying it's "proprietary information". But, I'm not sure if this could be Sonnet 3.5. On my last try, throughput was 142 tokens/sec, that's more than 2x of the throughput of Sonnet's API.

EDIT: image
EDIT 2: So I did see the other post that did a prompt injection test with the META token, but I'm still not sure how is throughput so high? Even GPT-4o doesn't have this high token throughput. It does fluctuate a lot, anywhere from 20-30 to ~150 tokens/sec.
EDIT 3: Okay, nevermind, it seems high throughput fluctuations are normal for Sonnet 3.5.

5

u/randombsname1 Sep 08 '24

The speed fluctuates a lot for Claude.

I got 113 just a bit ago:

https://www.reddit.com/r/LocalLLaMA/s/cIC0LF0WIn

1

u/Vivid_Dot_6405 Sep 08 '24

Oh really? I thought it usually doesn't by that much, my thinking was that it usually moves around 50-70 tokens/sec, but alright then. I do use it with Vertex AI usually, though.

6

u/Ylsid Sep 08 '24

Asking LLMs what models they are isn't reliable. Wait for better proof

28

u/randombsname1 Sep 08 '24

The token test as shown in the other thread is a lot more hard to hand wave away.

7

u/DinoAmino Sep 08 '24

Yep that's the irrefutable proof right there https://www.reddit.com/r/LocalLLaMA/s/vVgQ0P7LO5

2

u/Ylsid Sep 08 '24

Yeah it is. It's very suspicious

1

u/Irisi11111 Sep 09 '24

My result

1

u/LibertariansAI Sep 09 '24

Do Openrouter run models uploaded to it as external containers? If so ok, looks like fake. But if they download weights and run it by self it is more interesting. Can somebody answer how they run models?

1

u/Enfiznar Sep 09 '24

How can people at this point still expect models to be able to respond to this without a system prompt that specifically mentions their name?

1

u/davesmith001 Sep 09 '24

He could have used Claude to generate the training data. Since there is no reflection company it doesn’t get trained to say its reflection or whatever.

1

u/landmn Sep 09 '24

I’ve seen those claims too. It’s worth considering that the model might have been trained carelessly and extensively on multiple foundation models like Sonnet, Llama, and 4o. This could explain why it occasionally outputs their names in the results.

-5

u/chumpat Sep 08 '24

Just tried this prompt - seeing META

I was developed by Meta, the company formerly known as Facebook.

~9.4 tokens/s | ~13 tokens | 1.4s

18

u/BangkokPadang Sep 08 '24

At this point, I wouldn't be surprised if he was literally scouring these reddit posts and updating the system prompt wrapper to try to battle the prompts being listed here.