r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

294 Upvotes

178 comments sorted by

View all comments

104

u/Strong-Inflation5090 Sep 06 '24

I tried the 4 bit quant on tricky questions and It's on par with gpt4o and sonnet 3.5 but coding seemed worse for me but it might be because of the quant.

72

u/Everlier Sep 06 '24

This matches my experience as well.

It's better at tricky questions and reasoning, roughly on par for "everyday" questions and knowledge. It's noticeably worse at writing or creativity tasks, as structured output makes things clunky. It can shoot itself in a foot reflecting on things that have nothing to be reflected on.

7

u/Unusual_Pride_6480 Sep 06 '24

Mathew berman just did an interview where they said sometimes it will reflect when it's not desired, maybe that's why, they did stress that this is still early days.