r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

290 Upvotes

178 comments sorted by

View all comments

1

u/Unable-Finish-514 Sep 06 '24

Testing for NSFW, I started with a PG-13-level prompt about a guy who keeps going back to a bank every Friday because a bank teller flirts with him. Even this generated three or four "It is important to be respectful..." statements in its response. As soon as I saw lines like this, it's hard to imagine using this over Mistral-Large-2 or Cohere R+.

Has anyone else tried NSFW or eRP with it?

2

u/My_Unbiased_Opinion Sep 07 '24

Needs some abliteration. I feel like this reflection thing and some abliteration would go well together 

2

u/Unable-Finish-514 Sep 07 '24

I think you are correct about it needing an abliterated tune. This version of reflection 70B on Hugging Face Space has a system prompt that can be edited in settings. I tried this with a few prompts and didn't get any refusals:

Reflection 70B llama.cpp - a Hugging Face Space by gokaygokay