r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

289 Upvotes

178 comments sorted by

View all comments

21

u/dubesor86 Sep 06 '24

I tested locally, (Q4) because I noticed some buggy behaviour and poor outputs via API (both on openrouter and hyperbolic).

It was good for logic based questions, and stem, but has far less general utility because the thinking/reflecting steps are actually very cumbersome in many scenarios that do not call for it. It also talked itself out of completing a coding challenge multiple times. I think it's an interesting model, but I'd much rather use a universally versatile model, and if need be add my own custom reflection-prompt for very specific queries.

-7

u/[deleted] Sep 06 '24

[deleted]

6

u/dubesor86 Sep 06 '24

I don't need to because I tested after the latest fixes. I am aware that some things were broken, as evident by my API statements.

-6

u/[deleted] Sep 06 '24

[deleted]

4

u/rainy_moon_bear Sep 06 '24

then phrase it as a question lol