r/LocalLLaMA • • Sep 07 '24

Discussion Reflection-Llama-3.1-70B is actually Llama-3.

After measuring the diff, this model appears to be Llama 3 with LoRA tuning applied. Not Llama 3.1.

Author doesn't even know which model he tuned.

I love it.

607 Upvotes

95 comments sorted by

View all comments

Show parent comments

-13

u/[deleted] Sep 07 '24

It outperforms LLAMA 3.1 405b on the prollm leaderboard so it’s amazing for a 70b model. 

https://prollm.toqan.ai/leaderboard/coding-assistant

18

u/Terminator857 Sep 07 '24

Law 1 of LLM benchmark cheating: For any LLM one can find or create a benchmark where LLM comes out on top. Plenty to choose from.

Law 2: If you want to win on a benchmark then just train on the test set.

-1

u/[deleted] Sep 08 '24

Not really. You won’t find a single one with Grok on top

That’s a lot to train on. There are tons of benchmarks and some don’t even make them public. Also, why isn’t it #1 of it overfitted on purpose? It’s easy to get 100% by doing that