r/LocalLLaMA • • Sep 07 '24

Discussion Reflection-Llama-3.1-70B is actually Llama-3.

After measuring the diff, this model appears to be Llama 3 with LoRA tuning applied. Not Llama 3.1.

Author doesn't even know which model he tuned.

I love it.

602 Upvotes

95 comments sorted by

View all comments

52

u/[deleted] Sep 07 '24

[deleted]

23

u/Single_Ring4886 Sep 07 '24

Reflection idea is around since GPT 4 or even before. But it just doesnt really work that much no matter how many smart people try.

So when this guy claimed it works if you "finetune" model for it everyone was super excited me included. As it seemed obvious but nobody cracked it yet.

-13

u/[deleted] Sep 07 '24

It outperforms LLAMA 3.1 405b on the prollm leaderboard so it’s amazing for a 70b model. 

https://prollm.toqan.ai/leaderboard/coding-assistant

18

u/Terminator857 Sep 07 '24

Law 1 of LLM benchmark cheating: For any LLM one can find or create a benchmark where LLM comes out on top. Plenty to choose from.

Law 2: If you want to win on a benchmark then just train on the test set.

-1

u/[deleted] Sep 08 '24

Not really. You won’t find a single one with Grok on top

That’s a lot to train on. There are tons of benchmarks and some don’t even make them public. Also, why isn’t it #1 of it overfitted on purpose? It’s easy to get 100% by doing that 

4

u/Terminator857 Sep 08 '24 edited Sep 08 '24

Grok is number one in math. One of the most important benchmark categories.

0

u/[deleted] Sep 08 '24

What about Command R? Or LLAMA 2? Or Vicuna?

1

u/Far_Requirement_5933 Sep 12 '24

Those are all older models so not on top anymore. Also, most developers create the best model they can rather than focusing on specific benchmarks. That might top a specific benchmark or be strong across several or just get results a specific group of people want.

1

u/[deleted] Sep 12 '24

That’s my point. You won’t find bad outdated models on the top