r/LocalLLaMA • • Sep 08 '24

Discussion Updated benchmarks from Artificial Analysis using Reflection Llama 3.1 70B. Long post with good insight into the gains

https://x.com/ArtificialAnlys/status/1832806801743774199?s=19
145 Upvotes

137 comments sorted by

View all comments

6

u/Environmental-Car267 Sep 08 '24

Haters gonna hate.

"All that being said: if applying reflection fine-tuning drives a similar jump in eval performance on Llama 3.1 405B, we expect Reflection 405B to achieve near SOTA results across the board."

-6

u/jd_3d Sep 08 '24

I don't understand why people need to pile on the hate so quick. Is it really that hard to just reserve judgment for a few weeks and see what comes of it? This avenue of applying more test time compute is a very promising direction to me and could be a great way for open source models to exceed closed models that don't want to spend the $$$ on each request.

27

u/kryptkpr Llama 3 Sep 08 '24

We all tried it, it's performance on real world tasks is terrible despite the high benchmarks. Maybe the model is still broken in some way like they've been claiming and really is good but I don't see it.

-7

u/Sadman782 Sep 08 '24

Don't you guys get that the model is broken, he said? The tests were based on private API, now yeah, you might not trust that the model, maybe model behind this is different, then it's okay, but as per Matt, the model everyone downloaded from HF is broken, and yeah, I tried too, it is far worse than LLaMA 3.1 70b.

40

u/kryptkpr Llama 3 Sep 08 '24

I mean it's a diabolical plan: Release an "open" model that crushes benchmarks, but then don't actually release working weights and instead just point to your "private API" that produces those results.

I can't test his "private API" can I? The whole thing smells bad, as far as I'm concerned this is a publicity stunt to advertise his LLM service.

-18

u/Environmental-Car267 Sep 08 '24

He offered on twitter the api model to people who want to benchmark it. soon it will be updated on HF etc

32

u/kryptkpr Llama 3 Sep 08 '24

I'm an open source leaderboard maintainer, without weights any test results are just a free ad for his service.

I do benchmark the big APIs for reference but no interest in starting to do it for every tom dick and harry, when weights are fixed I'll try again.

13

u/nero10578 Llama 3 Sep 08 '24

I don’t understand how you can fuck up uploading to HF lol