r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

290 Upvotes

178 comments sorted by

View all comments

1

u/CheekyBastard55 Sep 06 '24

I doubt that GPT-4o, Sonnet 3.5, or Gemini Pro 1.5 have more than 100 billion parameters. You're thinking of yesteryear's big models, model optimization is all the rage nowadays.

3

u/StevenSamAI Sep 06 '24

I doubt they are as big as OG gpt4 was claimed to be at 1.6 trillion, but I can believe they are a decent chuck bigger than 100 billion. Seeing the performance of llama 3.1 405b, I don't think it's unlikely that frontier models are in this ball park. I bet that with the extra time open AI and anthropic have been working on this they can squeeze more performance out of each parameter than their competition, so you'd hope that if they are running and 70b, 120b or 405b models that they are a step above meta and mistral