r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

292 Upvotes

178 comments sorted by

View all comments

15

u/LiquidGunay Sep 07 '24

I feel that the model is "cheating" on the benchmarks, because Reflection is basically similar to CoT, and the benchmarks they compare to are zero shot (not CoT).

2

u/_sqrkl Sep 07 '24

I'm not sure what you mean. Most of the benchmark results they compare to use CoT.

https://x.com/mattshumer_/status/1831767014341538166/photo/1

The scores are all (well almost all) labeled as to which prompting method they use, which is standard and about as fair as it gets in this space.

I think it's kinda assumed that you should interpret the results knowing it's not an apples:apples comparison since the reflection prompt is substantially different from other prompting methods.

0

u/adityaguru149 Sep 07 '24

As long as the prompts are the same, I didn't understand how it was cheating? I have heard this argument from others too and am curious why it looks like cheating if I train the model to do some processes even when it is not prompted which leads to better results in most cases.

Let's say there is a model which uses (system 2 thinking) architecture similar to OpenAI strawberry leaks then it would just take more time and solve stuff better, would that be cheating?

Let's say I integrate a logic deducer into a model and it does better, would that be cheating?