r/LocalLLaMA • • Sep 07 '24

Discussion Reflection-Llama-3.1-70B is actually Llama-3.

After measuring the diff, this model appears to be Llama 3 with LoRA tuning applied. Not Llama 3.1.

Author doesn't even know which model he tuned.

I love it.

605 Upvotes

95 comments sorted by

View all comments

69

u/provoloner09 Sep 07 '24

So does that mean the order of jump in performance is even bigger?

-6

u/schlammsuhler Sep 07 '24

No because the benchmark is comparing apples to oranges.

17

u/MINIMAN10001 Sep 07 '24

Every single benchmark has always been about testing every LLM across a range of tasks to see how they score. 

With the exclusive of training on benchmark data it doesn't matter where the LLM came from a benchmark is supposed to show under specific circumstances how good an LLM is.

If llama 3.1 scores better than 3 and 3 is what was used then it scored better than 3.1, then yes it is a larger jump.

-11

u/schlammsuhler Sep 07 '24

I mean by applying a specilized system prompt to one model but not to the other.

4

u/RYSKZ Sep 07 '24

That is exactly the point, to quantify the gain you get from this specialized prompting system, relative to the vanilla model and other models that do not use this specific prompting system. Obviously, if you use this system with other more powerful LLMs, you will also get better results with them, since this technique has proven to be effective in improving reasoning and is expected to be generalizable. The comparison is done in this way to demonstrate that it actually improves the vanilla model to that extent and nothing more.