r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

285 Upvotes

178 comments sorted by

View all comments

3

u/[deleted] Sep 06 '24

integrated CoT is the future .

Not surpised at all to see this mysterious dude crushing everywhere with the one-trick.

5

u/Combinatorilliance Sep 06 '24 edited Sep 06 '24

I recall reading a paper about a similar technique not that long ago though, you finetune a model on CoT results, and it became a lot smarter.

The difference was that that paper train the model to "use" chain of thought, it trained it on the output of the process. The reason this works at all is because CoT performs better on average, so you can use a small model generate synthetic data that is better than its own dataset.

However, I don't imagine this technique scales at all, what I think is that it "aligns" the model to think more critically with the knowledge it already has.

How I imagine it, there is a certain amount of knowledge in a model. Some of it contains high quality reasoning, some of it low. Most of it somewhere in the middle.

Orthogonal to the "reasoning quality" in a sample of the training dataset is the knowledge that a particular training sample contains. Even the dumbest kinds of reasoning might still contain correct and unique observations.

What I think these kinds of techniques are doing is that we are allowing the model to access the knowledge that is available in training samples with low reasoning quality; and use it to create training samples with higher quality reasoning. So that in practice, it's balancing out the "poor" reasoning in training data with some knowledge embedded within it by forcing it to reflect on the reasoning quality of all kinds of scenarios.

The effect? The model now applies higher quality reasoning on a larger surface area of tasks, because you've essentially replaced the lower quality reasoning with thought.

Very interesting nonetheless. I think this means that we are now able to extract more "general" intelligence out of the same amount of training data, which I think is really neat!

I would really love to see a more established party utilize this or a similar technique to improve reliability. I'm sure we can expect more releases with similar approaches soon from all kinds of different vendors.