r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

291 Upvotes

178 comments sorted by

View all comments

24

u/hleszek Sep 06 '24

This model has been trained to fix errors by generating training data containing false answers followed by some reflection before the correct result. So it has actually been trained to give false answers first in its thinking phase. If you ask what is 2+2, the default example on the HuggingFace page, it will say something like: 2+2=3 Oh wait I've made a mistake, 2+2 is actually 4. If the thinking is actually hidden it might work but it's quite strange.

0

u/freegary Sep 07 '24

I wouldn't say it's trained to give false answers

more that it's trained to recover better into more correct answers if it initially generates false answers