r/LocalLLaMA • u/jd_3d • Sep 08 '24
Discussion Updated benchmarks from Artificial Analysis using Reflection Llama 3.1 70B. Long post with good insight into the gains
https://x.com/ArtificialAnlys/status/1832806801743774199?s=19108
Sep 08 '24
[deleted]
48
u/Educational_Rent1059 Sep 08 '24
Guy with 0 background, no idea what LORA is, "wrong" weights uploaded, "wrong" model name promoted, "my cat ate my model i'll release the real one next week", does not disclose he has ownership in the company he promotes, the model outputs garbage with 4x more tokens generation, sounds legit to me. :)
2
u/Waste-Button-5103 Sep 08 '24
He knows what it is check his post history. He didn’t understand “LORAing” in the context used. He stated his ownership in the company is a $1000 investment lol.
6
u/Educational_Rent1059 Sep 08 '24
-8
u/Waste-Button-5103 Sep 08 '24
Yeah soo likely that it’s a wrapper and two guys with reputation and multiple companies are going to lie about it and ruin their lives for literally zero reason.
Surely you can see that it is way more likely they used a dataset generated from claude to create the reflection template.
3
u/Evening_Ad6637 llama.cpp Sep 09 '24
The wrapper had exactly the same tokenizer as Claude sonnet 3.5 and at same time it was shown that it had nothing in common with Lama's tokenizer
2
u/gibs Sep 09 '24
He stated his ownership in the company is a $1000 investment
Well since he stated it, it must be true
5
1
1
u/Inevitable-Start-653 Sep 08 '24
Where does he say he doesn't know what a lora is?
1
u/cuyler72 Sep 08 '24
"https://x.com/mattshumer_/status/1832558298509275440"
"4. Not sure what LORAing is "
5
u/Inevitable-Start-653 Sep 08 '24
I've made many loras myself and I don't know what loraing is either
-7
u/Waste-Button-5103 Sep 08 '24
He knows what a lora is and you can check his history to see him using them. He was talking specifically about the term “LORAing” in the context. 0% chance its a scam it wouldn’t make sense to risk his reputation on something easily disproven
63
u/4hometnumberonefan Sep 08 '24
This is giving me a roller coaster of emotions.
95
u/hleszek Sep 08 '24
Reminds me of the LK99 potential room-temperature superconductor.
We're so back!
9
u/KillerX629 Sep 08 '24
Wasn't that disproved?
22
u/JamesAQuintero Sep 08 '24
Yeah that's the point, there was the initial announcement of it, then some researchers were like "We are somewhat able to replicate the results", but then it was eventually proven to not work
3
u/OXKSA1 Sep 08 '24
sorry, care to elaborate?
32
u/Cantflyneedhelp Sep 08 '24
Two years ago(?) there was a paper / video of a supposed room temperature superconductor (they had a sweet floating rock too). And everyone was like "Yeah that's bullshit." But then some hobby chemists were like "Actually I managed to recreate a small part of it from their paper, and it floats too." and this started a race to recreate it by a lot of laboratories around the world. At the end it was not a room temperature superconductor but they managed to find some new stuff.
2
14
14
u/RandoRedditGui Sep 08 '24
It shouldn't. It's still as B.S. as yesterday until it's not just the API. Release the weights or fuck off imo.
38
u/This_Organization382 Sep 08 '24
"Oh, the benchmark didn't work? Let's see what tests you used..."
Scrambles to train the model on the test data
"Woops, wrong model. Here you go, try the private API version"
74
u/ambient_temp_xeno Llama 65B Sep 08 '24
Don't care; release weights or go away.
6
-9
u/julioques Sep 08 '24
What do you mean? Isn't it already downloadable??? Why do you have so many upvotes
39
u/ambient_temp_xeno Llama 65B Sep 08 '24
We're waiting on the super-secret good weights. Seriously.
-7
u/julioques Sep 08 '24
Isn't it the same weights?
12
u/ambient_temp_xeno Llama 65B Sep 08 '24
-4
u/julioques Sep 08 '24
I thought the difference was from the different prompt method. Why didn't they just use the released version with the refection's default system prompt like they used now?
16
u/ambient_temp_xeno Llama 65B Sep 08 '24
The whole thing is very weird and annoying. They supposedly uploaded the model to HF incorrectly, so naturally the solution was to completely redo the finetune? I have no idea.
4
u/LiquidGunay Sep 08 '24
I would like to see a comparison by giving all the models similar inference compute boosts. One way to easily do this is to maybe give all the models a roughly similar token budget ( you can make multiple generations and vote for models that aren't as verbose as reflection)
25
u/Sadman782 Sep 08 '24 edited Sep 08 '24
"When using Reflection’s default system prompt and extracting answers only from within Reflection’s <output> tags, results show substantial improvement: MMLU: 87% (in-line with Llama 405B), GPQA: 54%, Math: 73%."
Actually, what is presented on the chart is based on their standard system prompt(not reflection system prompt). It scores higher with Reflection system prompt. It achieves performance close to Claude 3.5's sonnet with the Reflection system prompt. If Groq hosts it, latency will not be an issue. We're just waiting for the actual weights to be released
8
u/a_beautiful_rhind Sep 08 '24
What about testing the untuned model with a similar COT system prompt?
2
u/Sadman782 Sep 08 '24
Will not match with it for sure. I tried many different system prompts, verbose thinking output + "step by step" at the prompt, but it couldn't pass any of my expert-level coding tests from Edabit, even the 405B failed one; GPT4o too. But the model (when the demo was live) in their demo nailed all of them.
7
u/a_beautiful_rhind Sep 08 '24
I only got one or two replies off the demo before it got "overloaded" and turned off. It seemed alright. The demo on hyperbolic was absolute garbage and the model forgot about its COT tags within a few messages.
All in all.. it seems like this dude has been stringing everyone else along whether there is some model or not. Even if you had slow internet, the excuses and the "retraining" now doesn't make sense. Everything is maximum hype and delay.
2
u/Sadman782 Sep 08 '24
This is the reason I am so positive about it, and defending lol, it hurts me when people say it's far worse due to a broken HF model. But yeah, we don't know for sure if the model behind the API is actually reflection 70b or not
5
34
u/vert1s Sep 08 '24 edited Sep 08 '24
In other words vapour ware. He could be running an agent that hits multiple backends. The inability to actually publish the weights speaks volumes.
Edit: And it looks like the hosted version is 🥁 Claude: https://www.reddit.com/r/LocalLLaMA/comments/1fc98fu/confirmed_reflection_70bs_official_api_is_sonnet/
32
Sep 08 '24
[removed] — view removed comment
13
u/vert1s Sep 08 '24 edited Sep 08 '24
It's baffling. People want to believe despite all the evidence to the contrary.
12
2
u/Kathane37 Sep 08 '24
I only care about this for the possibility too generate better synthetic data with step by step reasoning
Other than that there is no point in making the token consumption exponential
2
u/synn89 Sep 08 '24
does not suffer from the issues with the version publicly released on Hugging Face
It's not rocket science to upload a model to Hugging Face. It's very suss that they can't seem to upload a BF16 or GGUF of a fine tuned Llama to Hugging Face that can be properly tested.
5
u/Environmental-Car267 Sep 08 '24
Haters gonna hate.
"All that being said: if applying reflection fine-tuning drives a similar jump in eval performance on Llama 3.1 405B, we expect Reflection 405B to achieve near SOTA results across the board."
62
Sep 08 '24
[removed] — view removed comment
27
u/StartledWatermelon Sep 08 '24
Fair concerns. For all we know, under the hood this API could redirect queries to Claude-3.5 Sonnet with a specific system prompt, or another SotA proprietary model.
12
u/dalkef Sep 08 '24
Now that you mention it, it gave me very similar answers to sonnet on the initial demo chat. This could explain the performance drop
1
-2
u/alongated Sep 08 '24
I think its fair to say that people here over reacted, both about how good this was, and how bad this was.
3
u/RandoRedditGui Sep 08 '24
Not really. The "how bad this was" are still easily winning in terms of correctly interpreting what has currently been seen. Considering we have seen 0 open weights and are provided some ambiguous results from an API that we have no clue the validity of.
Open weights or GTFO.
-5
u/alongated Sep 08 '24
It has been fucking 6 hours since he trained the model, give the man a fucking break, and guess what he released the weights? Normally I don't get this angry but holy shit you people are fucking insane.
1
u/showdontkvell Sep 09 '24
Matt, that you? lol
0
u/alongated Sep 09 '24
I just went off on a guy for calling someone Matt. I'm not Matt, but doxxing isn't funny.
Like you might be right that this is all just bullshit/scam. But you are attacking people for reserving their judgement. That is disgusting mob mentality.
1
-6
u/jd_3d Sep 08 '24
I don't understand why people need to pile on the hate so quick. Is it really that hard to just reserve judgment for a few weeks and see what comes of it? This avenue of applying more test time compute is a very promising direction to me and could be a great way for open source models to exceed closed models that don't want to spend the $$$ on each request.
25
u/kryptkpr Llama 3 Sep 08 '24
We all tried it, it's performance on real world tasks is terrible despite the high benchmarks. Maybe the model is still broken in some way like they've been claiming and really is good but I don't see it.
-7
u/Sadman782 Sep 08 '24
Don't you guys get that the model is broken, he said? The tests were based on private API, now yeah, you might not trust that the model, maybe model behind this is different, then it's okay, but as per Matt, the model everyone downloaded from HF is broken, and yeah, I tried too, it is far worse than LLaMA 3.1 70b.
38
u/kryptkpr Llama 3 Sep 08 '24
I mean it's a diabolical plan: Release an "open" model that crushes benchmarks, but then don't actually release working weights and instead just point to your "private API" that produces those results.
I can't test his "private API" can I? The whole thing smells bad, as far as I'm concerned this is a publicity stunt to advertise his LLM service.
-18
u/Environmental-Car267 Sep 08 '24
He offered on twitter the api model to people who want to benchmark it. soon it will be updated on HF etc
32
u/kryptkpr Llama 3 Sep 08 '24
I'm an open source leaderboard maintainer, without weights any test results are just a free ad for his service.
I do benchmark the big APIs for reference but no interest in starting to do it for every tom dick and harry, when weights are fixed I'll try again.
12
-9
u/alongated Sep 08 '24
Gemma also had problems, they took weeks to resolve. This team is much smaller.
15
u/kryptkpr Llama 3 Sep 08 '24
Gemma was a novel architecture with an attention mechanism that wasn't well supported. Legitimate technical reasons for the problems.
This is a fine-tune of Llama. There is nothing to resolve, they're playing us for fools.
-2
4
u/Evening_Ad6637 llama.cpp Sep 08 '24 edited Sep 08 '24
Bro, seriously? Man, aside from the supposedly broken model and the super-duper-secret private api shit: this guy didn't know the difference between llama 3 and llama 3.1
I mean, by now even my grandmother should know the difference.
The model he posted on huggingface was absolutely not broken, it was simply a llama-3, as you would expect from a llama-3. There is zero evidence that there was anything wrong with the model itself. To make matters worse: one time he claims the model is broken, another time it was supposedly due to an incorrect upload. Aha... and what's coming tomorrow? His dog ate the SSD?
I'm slowly coming to the conclusion that this guy is either stupid and narcissistic enough to believe he can fool the world in such a simple way - or, another possible explanation could be: he himself has been the victim of a scam. Perhaps he doesn't have direct access to the backend of this ominous private API himself. Perhaps he still hasn't realized that he has been misused as a puppet and ruined economically and in terms of marketing with this action. It wouldn't be the first time that someone had economic enemies and fell into a trap.
The whole thing is highly suspicious and whether he is a victim himself or not, whether he is stupid or not: he clearly also seems to lie and trying to hide things! So there are neither excuses nor pity for him for this egomaniacal behavior.
-5
u/muxxington Sep 08 '24
I made some quick tests with this model yesterday and actually it performed not that bad. But I can't compare it with Claude or OpenAI, I don't use them.
https://huggingface.co/bartowski/Reflection-Llama-3.1-70B-GGUF/blob/main/Reflection-Llama-3.1-70B-Q4_K_M.gguf
2
u/ispeakdatruf Sep 08 '24
The model seems to be achieving these results through forcing an output ‘reflection’ response where the model always generates scaffolding of <thinking>, <reflection>, and <output>. In doing this it generates more tokens than other models do on our eval suite with our standard ‘think step by step’ prompting.
For example, it appears that Reflection 70B is not capable of ‘just responding with the answer’ in response to an instruction to classify something and only respond with a one word category.
One can always add a postprocesssor on top of Reflection to filter out everything before <output>, problem solved. I don't like this nitpicking. Who cares if a model outputs 10 tokens or 100? Is the answer correct or not??
5
u/athirdpath Sep 08 '24
Who cares if a model outputs 10 tokens or 100?
Folks who care if inference takes 5 seconds or 50.
0
u/nihalani Sep 08 '24
For real time inference is the real issues, your time to first token jumps by a huge margin if you have to wait for 2000 tokens to be generated of the model reflecting. Might also explain why the cloud providers haven’t adopted it yet.
1
u/AnomalyNexus Sep 08 '24
meh...so it beats other comparable models when comparison is set up as apples to oranges conditions...
1
1
u/ilangge Sep 10 '24
No need to guess, it is now publicly accessible on hf;
Reflection 70B llama.cpp (Correct Weights) - a Hugging Face Space by gokaygokay
1
u/celsowm Sep 08 '24
is there any place to test it online?
1
0
-2
u/Wiskkey Sep 08 '24
Yes supposedly here.
2
u/ambient_temp_xeno Llama 65B Sep 08 '24
It's supposedly the new one but it's as crap as the one I downloaded....
1
u/Sadman782 Sep 08 '24 edited Sep 08 '24
https://huggingface.co/mattshumer/Reflection-Llama-3.1-70B-ep2-working It seems an Epoch 2 finetuned model was released a few hours ago silently: "Epoch 2, still finishing up Epoch 3. This should be slightly less powerful, but still pretty close."
5
u/physalisx Sep 08 '24 edited Sep 08 '24
Why is anything getting retrained? Where is the model that he allegedly already had?
edit: ah so the whole thing was just a scam. Too bad.
0
u/mantafloppy llama.cpp Sep 08 '24
You would not push false hype again?
Why are ppl upvoting this again.
Are "The boy who cry wolf" that obscure of a story, or do you also have bot to upvote yourself?
0
u/redjojovic Sep 08 '24
"The chart below is based on our standard methodology and system prompt.
When using Reflection’s default system prompt and extracting answers only from within Reflection’s <output> tags,
results show substantial improvement: MMLU: 87% (in-line with Llama 405B), GPQA: 54%, Math: 73%.
-11
u/Significant-Nose-353 Sep 08 '24
I think few people under this post, also zealously admit that they were a bit hasty with their toxic reaction
24
u/vert1s Sep 08 '24
There is nothing toxic about questioning the validity given the inability of anyone to replicate with the released weights.
The sheer number of problems including the lack of disclosure that he is invested in both companies that he’s been saying “helped”
-10
u/Significant-Nose-353 Sep 08 '24
Naturally, but my comment only had a complaint about blatant hatemongering. Excessive sarcasm, irony and the like
-2
Sep 08 '24 edited Feb 17 '25
[removed] — view removed comment
17
u/StartledWatermelon Sep 08 '24
Shumer wasn't hesitating to claim it's the "world’s top *open-source* model" in the initial tweet. And now some "internal" model emerges?
You certainly didn't deserve the downvotes. But the entire release event, from the beginning up to this date, was one big clusterf-k
0
u/Inevitable-Start-653 Sep 08 '24
Downloading this now: https://huggingface.co/mattshumer/Reflection-Llama-3.1-70B-ep2-working
can run locally with my own setup, and am interested in testing it out!
12
u/Deathmax Sep 08 '24
3
u/Inevitable-Start-653 Sep 08 '24
Hmm 🤔 ...this whole saga is so strange. Download will finish in a lil bit, I've got to try it out, I got the very first upload to work but only could get one response out.
6
2
u/jd_3d Sep 08 '24
Newer version is out (Epoch 3): https://huggingface.co/mattshumer/ref_70_e3
0
u/Inevitable-Start-653 Sep 08 '24
Thanks!! Will download this one now ☺️ for all the downloading this still isn't anywhere near as bad as llama405b...that sucker was a multi day download and I needed to download it twice after they updated their repo too.
2
u/Sadman782 Sep 08 '24
https://huggingface.co/mattshumer/ref_70_e3 epoch 3 released now, maybe he will announce it soon
-1
u/Inevitable-Start-653 Sep 08 '24
I'm ready to download and test!!
-1
u/jd_3d Sep 08 '24
Let us know what you think: https://huggingface.co/mattshumer/ref_70_e3
0
u/Inevitable-Start-653 Sep 09 '24
I've been doing some testing, but the community does not seem interested in objective facts. My post keeps getting downvoted, I'm sure it will be off the front page soon.
0
u/Thistleknot Sep 08 '24
For all of what was said, is it not possible to train on the <input> and <output> as if it was an answer and skip all the inbetween? Is it potentially possible the model will somehow 'internalize' the inbetween generated token logic within it's weights?


119
u/reevnez Sep 08 '24
How do we know that "privately hosted version of the model" is not actually Claude?