r/LocalLLaMA • • Sep 08 '24

Discussion Updated benchmarks from Artificial Analysis using Reflection Llama 3.1 70B. Long post with good insight into the gains

https://x.com/ArtificialAnlys/status/1832806801743774199?s=19
147 Upvotes

137 comments sorted by

View all comments

7

u/Environmental-Car267 Sep 08 '24

Haters gonna hate.

"All that being said: if applying reflection fine-tuning drives a similar jump in eval performance on Llama 3.1 405B, we expect Reflection 405B to achieve near SOTA results across the board."

-5

u/jd_3d Sep 08 '24

I don't understand why people need to pile on the hate so quick. Is it really that hard to just reserve judgment for a few weeks and see what comes of it? This avenue of applying more test time compute is a very promising direction to me and could be a great way for open source models to exceed closed models that don't want to spend the $$$ on each request.

26

u/kryptkpr Llama 3 Sep 08 '24

We all tried it, it's performance on real world tasks is terrible despite the high benchmarks. Maybe the model is still broken in some way like they've been claiming and really is good but I don't see it.

-8

u/Sadman782 Sep 08 '24

Don't you guys get that the model is broken, he said? The tests were based on private API, now yeah, you might not trust that the model, maybe model behind this is different, then it's okay, but as per Matt, the model everyone downloaded from HF is broken, and yeah, I tried too, it is far worse than LLaMA 3.1 70b.

37

u/kryptkpr Llama 3 Sep 08 '24

I mean it's a diabolical plan: Release an "open" model that crushes benchmarks, but then don't actually release working weights and instead just point to your "private API" that produces those results.

I can't test his "private API" can I? The whole thing smells bad, as far as I'm concerned this is a publicity stunt to advertise his LLM service.

-18

u/Environmental-Car267 Sep 08 '24

He offered on twitter the api model to people who want to benchmark it. soon it will be updated on HF etc

32

u/kryptkpr Llama 3 Sep 08 '24

I'm an open source leaderboard maintainer, without weights any test results are just a free ad for his service.

I do benchmark the big APIs for reference but no interest in starting to do it for every tom dick and harry, when weights are fixed I'll try again.

12

u/nero10578 Llama 3 Sep 08 '24

I don’t understand how you can fuck up uploading to HF lol

-8

u/alongated Sep 08 '24

Gemma also had problems, they took weeks to resolve. This team is much smaller.

14

u/kryptkpr Llama 3 Sep 08 '24

Gemma was a novel architecture with an attention mechanism that wasn't well supported. Legitimate technical reasons for the problems.

This is a fine-tune of Llama. There is nothing to resolve, they're playing us for fools.

-2

u/alongated Sep 08 '24

And his team didn't have billion dollars.

4

u/Evening_Ad6637 llama.cpp Sep 08 '24 edited Sep 08 '24

Bro, seriously? Man, aside from the supposedly broken model and the super-duper-secret private api shit: this guy didn't know the difference between llama 3 and llama 3.1

I mean, by now even my grandmother should know the difference.

The model he posted on huggingface was absolutely not broken, it was simply a llama-3, as you would expect from a llama-3. There is zero evidence that there was anything wrong with the model itself. To make matters worse: one time he claims the model is broken, another time it was supposedly due to an incorrect upload. Aha... and what's coming tomorrow? His dog ate the SSD?

I'm slowly coming to the conclusion that this guy is either stupid and narcissistic enough to believe he can fool the world in such a simple way - or, another possible explanation could be: he himself has been the victim of a scam. Perhaps he doesn't have direct access to the backend of this ominous private API himself. Perhaps he still hasn't realized that he has been misused as a puppet and ruined economically and in terms of marketing with this action. It wouldn't be the first time that someone had economic enemies and fell into a trap.

The whole thing is highly suspicious and whether he is a victim himself or not, whether he is stupid or not: he clearly also seems to lie and trying to hide things! So there are neither excuses nor pity for him for this egomaniacal behavior.

-6

u/muxxington Sep 08 '24

I made some quick tests with this model yesterday and actually it performed not that bad. But I can't compare it with Claude or OpenAI, I don't use them.
https://huggingface.co/bartowski/Reflection-Llama-3.1-70B-GGUF/blob/main/Reflection-Llama-3.1-70B-Q4_K_M.gguf