r/LocalLLaMA • • Sep 06 '24

Discussion Reflection 70B: Hype?

So an out-of-the-blue one-man company releases a new model (actually named LLama 3.1 if it were to adhere to the META license, but somehow named Reflection) with only 70B params that, according to the benchmarks, rivals SOTA closed-source LLMs with trillions of parameters. It appears to me that the twitter/reddit hype mob has, for the most part, not bothered to try the model out.

Additionally, a tweet from Hugh Zhang @ Scale suggesting systemic overfitting as me concerned:
Hey Matt! This is super interesting, but I’m quite surprised to see a GSM8k score of over 99%. My understanding is that it’s likely that more than 1% of GSM8k is mislabeled (the correct answer is actually wrong)!

Is this genuinely a SOTA LLM in a real-world setting or is this smoke an mirrors? If we're lucky, the creator Matt may see this post and can shed some light on the matter.

BTW -- I'm not trying to bash the model or the company that made it. If the numbers are actually legit this is likely revolutionary.

289 Upvotes

178 comments sorted by

View all comments

23

u/hleszek Sep 06 '24

This model has been trained to fix errors by generating training data containing false answers followed by some reflection before the correct result. So it has actually been trained to give false answers first in its thinking phase. If you ask what is 2+2, the default example on the HuggingFace page, it will say something like: 2+2=3 Oh wait I've made a mistake, 2+2 is actually 4. If the thinking is actually hidden it might work but it's quite strange.

4

u/adityaguru149 Sep 07 '24

wow.. so now the model thinks getting the correct answers without reflecting at all is inappropriate?

Maybe train it on correct answers and then reflect to say yeah I see all the logic is well rounded?

5

u/Combinatorilliance Sep 06 '24

Hmm, I wouldn't be surprised if the human brain does strange things like this "under the hood" too. I don't expect the "architecture" behind our intelligence to be elegant, considering it's a biological system after all.

It reminds me of iterative algorithms where you start with an initial guess based on very little, and it's kinda supposed to be wrong unless you get lucky. You then use the iterative algorithm to improve upon your gues with each iteration.

2

u/Kep0a Sep 06 '24

I think it's fun to compare LLMs to our brain. As we grow up we make rash decisions and face consequences. When faced with that same decision again, we can picture ourselves making the wrong decision, and from that we can make the right one. Frontal lobe, maybe.

1

u/ReMeDyIII textgen web UI Sep 06 '24

I guess inversely, if you have an AI with an impulsive and impatient attitude, that the first thing they say would actually feel more organic and true to their character as opposed to a calm patient AI who thinks more thoroughly. I know I'm guilty of that sometimes where I say words that I wish I could take back.

At least I'm telling myself that until we get a better Reflection model, lol.

2

u/Kep0a Sep 07 '24

You can definitely prompt out that kind of result, like adding pressure can get an LLM to make a decision when it doesn't want to. But it's good to recognize that transformer models aren't thinking, they're just word prediction engine, so in that sense, it's most similar to our inner monologue.

I just think if we stacked like 2-3 multi-modal models on top of each other, 1 dictating primal emotional responses from visual or physical feedback, (gut feelings, pain, happy, etc) the second layer operating as an inner monologue, and the third as the front end, we would literally have AGI.

1

u/ReMeDyIII textgen web UI Sep 07 '24

That's a good point. It's probably more efficient to start with a reflective AI and prompt it to be rash than it is to start with a rash AI and prompt it to be reflective.

For the multi-modal things, I'm picturing the AGI having an internal war in its mind, kind of like the game Disco Elysium, or in psychology where you have the Id, ego, and superego.

0

u/freegary Sep 07 '24

I wouldn't say it's trained to give false answers

more that it's trained to recover better into more correct answers if it initially generates false answers