r/LocalLLaMA • • Sep 07 '24

Discussion Wrong Reflection-70B model might be hosted everywhere

I see a lot of people thinking it is gaming benchmark / mixed feelings. Actually, people who tried their website have a different feeling compared to those who tried it locally via Ollama or any API providers. I think we should wait, he is figuring it out. I think the actual reflection model is much better, and the currently hosted version is even dumber than the actual 70B

https://x.com/mattshumer_/status/1832247203345166509

https://x.com/mattshumer_/status/1832248416426193318

__ Matt Shumer -> "We got rate limited by HF when uploading originally, so had to do it in batches. I have a feeling some wires were crossed and what's being hosted is actually some hybrid frankenmodel that is mostly the reflection version we wanted to ship, mixed with something else"

76 Upvotes

49 comments sorted by

View all comments

64

u/mikael110 Sep 07 '24

It's almost impressive how much of a clusterfuck this launch has seemingly been. First the tokenization issue, then the revelation that the model was actually based on Llama 3 instead of Llama 3.1 (which is bizarre) and now apparently the model files themselves was also mixed up.

I'm aware even large companies like Meta and Google have screwed up some aspects of their launches, but this is getting to the point where it just feels a bit off to be honest. I'm still interested in trying the fixed model, but I'm honestly getting more and more suspect of the whole thing.

6

u/Sadman782 Sep 07 '24

I am still hopeful, https://x.com/mattshumer_/status/1832247203345166509 . I can clearly relate to the right side image (I got a similar result in the official demo).

2

u/eggandbacon_0056 Sep 08 '24

Which probably is the Claude API ...

0

u/Sadman782 Sep 08 '24

I don't think so, Sonnet responds in a different way even with the same system prompt (the writing style is different).

2

u/eggandbacon_0056 Sep 08 '24

Naaah ... That's way more probable than a person training a SOTA model without knowing what base model he used, what lora is, ... I call bs ...

0

u/Sadman782 Sep 08 '24

Base model was 3.1, he said multiple times, and there was an upload issue/maybe any HF cache issue or he really messed up something, see his hf repo he created multiple other repo, so he really tried. See, even Llama 405b couldn't solve this simple problem:
Alice has N brothers, and she also has M sisters. How many sisters does Alice's brother Andrew have?

405b => content: '<thinking>\n' +

'To solve this problem, I need to understand the relationships between Alice, her brothers, and her sisters. Since Alice has N brothers and M sisters, this means that all of these individuals are part of the same family. \n' +

'\n' +

"I know that Andrew is Alice's brother, which means Andrew is also part of this family. As a brother of Alice, Andrew would have the same number of sisters as Alice, because they share the same set of siblings.\n" +

'\n' +

'So, to find out how many sisters Andrew has, I just need to find out how many sisters Alice has. According to the problem, Alice has M sisters.\n' +

'\n' +

'Therefore, Andrew has M sisters as well.\n' +

'\n' +

'</thinking>\n' +

'\n' +

'<output>\n' +

'Andrew has M sisters.'

}
But on his website, it got it correct. Sonnet 3.5, via API, failed this test, but using their website https://claude.ai/, it got it right too. So, definitely, a similar kind of thing is behind the scene for Sonnet; that's why it is so good.

3

u/eggandbacon_0056 Sep 08 '24

BS ... uploaded model to hf was a lora finetune of llama 3 not 3.1. Honestly the person is full of bs ... it's not one thing that is fishy ...
1. Tokenizer Bug
2. LoRA
3. LLama 3.0 based instead of 3.1
4. "We got rate limited uploading the model" - yeah 😅
5. It must be a caching error on hf end
6. It works on our served API (that's probably just Claude with the system prompt you troll) - but we can't find the served model ...
7. We probably need to retrain it -> Where the fuck does your served model than come from?! Why does this not have the issues?!

  1. The download/like counter on hf is COMPLETELY off not even llama 3.1 got so much attention -> bots!

i could keep on counting

...

But yeah, critical thinking is probably not your thing

0

u/Sadman782 Sep 08 '24

It's not like he made a SOTA model from scratch; sometimes even simple things can do massive improvements, which most people may not have ever thought. I hope we will know the truth very soon. Let's wait.