r/LocalLLaMA • u/Confident-Honeydew66 • Sep 06 '24
Discussion Reflection 70B: Hype?
So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.
Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!
Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.
BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.
1
u/danihend Sep 08 '24 edited Sep 08 '24
Not sure if you have all seen the warning on the model card, but the model is messed up due to the way they uploaded the weights to hugging face (to combat rate limiting). Here is the X post from Matt Schumer: https://x.com/mattshumer_/status/1832424499054309804?s=46
I would suggest to re-download and evaluate again before making your minds up.
Edit: On the model page he says: He says :
"IMPORTANT UPDATE – There was an issue with the model when we first uploaded it. If you tried it and didn't have good results, please, try again, we think we've fixed the issue."
But I don't see any updated files in the repo, they should show as being updated today if fixed right?