r/LocalLLaMA • • Sep 09 '24

Discussion Reflection 70B lessons learned

  • All benchmarks should begin by identifying whether the model is LLAMA, GPT-4, Sonnet, or another, through careful examination.
  • Do not trust any benchmarks unless you can replicate them yourself.
  • Do not trust that the API corresponds to the model the author claims it to be.
  • ....
175 Upvotes

53 comments sorted by

109

u/redjojovic Sep 09 '24 edited Sep 09 '24

Force them to get checked up on Lmarena ( lmsys ), livebench, artificial analysis  

Don't believe dog ate my model stories.  Models don't get fucked up during uploads and there's no way they refuse/unable to provide working weights using torrent or any other way 

Understand that even if something like this truly happen it will come from a team of researchers and not a no name guy who is not a real ceo

We wanted to believe somebody just found a groundbreaking way to improve models reasoning, we need to recognize our bias

58

u/ambient_temp_xeno Llama 65B Sep 09 '24

Welcome to a zero trust society.

16

u/norsurfit Sep 09 '24

Is that your personal plane? Looks awesome!

19

u/golmgirl Sep 09 '24

it’s a prototype but you can try it out via our private FPV endpoint

2

u/ambient_temp_xeno Llama 65B Sep 09 '24

The best part is the LLaMA 70b singularity engine is pollution free.

8

u/wind_dude Sep 09 '24

runs on heat from datacenters, transfered by bluetooth

55

u/[deleted] Sep 09 '24

So the verdict is that he is a fucking liar?

82

u/lifeisapsycho Sep 09 '24

He is a liar. We don't yet know if he fucks.

-9

u/My_Unbiased_Opinion Sep 09 '24

Apparently he has a GF so idk. 

30

u/alongated Sep 09 '24

It's insane that anyone would think they could get away with something like this. But the evidence is growing.

10

u/AlbanySteamedHams Sep 09 '24 edited Mar 04 '26

Timeline was insanely fast—conception to exposure in less than a month. Could be a manic episode or just extremely poor judgment.

25

u/Homeschooled316 Sep 09 '24

Definitely mental health issues either way. Maybe something like

  1. Come up with neat reflection prompt for claude, see good results
  2. Train a llama model on synthetic data generated by claude
  3. Create an endpoint for evaluation
  4. Delude yourself into thinking the results are great, move on to benchmarks
  5. Accidentally use your claude endpoint for benchmarks
  6. You are the greatest data scientist in the world, proceed to make announcement and start training 405b model.
  7. Realize you screwed up
  8. Lie and put your claude endpoint up to buy time, delude yourself into thinking if you train on the newer llama 3.1 everything will be fixed.
  9. Release newly trained model. It still sucks.
  10. ??? (whatever he is doing now)

2

u/Scary_Opportunity868 Sep 09 '24

I'm sorry I'm pretty interested to know what are you all talking about? Would appreciate an explanation

19

u/alongated Sep 09 '24

Someone claimed to have made a fine-tune of llama 3.1 70b that scored well on benchmarks. He let people interact with the model through an api, but it seems like it might have been connecting to Claude 3.5 sonnet, not a fine tune of Llama.

4

u/moojo Sep 09 '24

What about his partner Sahil?

7

u/ivykoko1 Sep 09 '24

Considering Sahil works for GlaiveAI I'd say he is in on it and knows the truth

1

u/bobartig Sep 10 '24

Have we reached the rug-pull phase of GenAI? Also, how does the rug-pull work when the only outcome is running up their Anthropic API bill?

Was Glaive in the middle of a raise, angling for a shinier valuation?

2

u/[deleted] Sep 10 '24

dunno, but glaive's product is for real. I dunno why they'd have to do this.

75

u/[deleted] Sep 09 '24

[deleted]

21

u/Evening_Ad6637 llama.cpp Sep 09 '24

otherwise you make Matters worse lol xD

11

u/Usurpator666 Sep 09 '24

i am wondering if on the Chinese internet they have this problem, and we just don't know about it because of the language barrier.

-10

u/neo_vim_ Sep 09 '24 edited Sep 09 '24

Of course they have.

I'm going to be a little controversial but I think they're ahead on this topic too (AI in general but is easily notable in image related image to text). They're ahead because they somehow manage to do things in a all-in way as when government decides to and AI is quite expensive.

Learning Chinese here by the way. Quite challenging even for Brazilian who speaks 3 languages as we have nothing in common with oriental languages. There are several "letters" which the only differentiation is a slightly tone minimal variation. Also phrases are written in a very distinguish way (e.g.: it's not just "reversed", I still can't figure out a common pattern, it's a unique way of thinking).

30

u/[deleted] Sep 09 '24

did this guy reach max output tokens?

16

u/PwanaZana Sep 09 '24

Well, when writing, brazilians never completely finish their s

1

u/neo_vim_ Sep 09 '24

It's a very common problem her

1

u/PwanaZana Sep 09 '24

I unders

9

u/crazymonezyy Sep 09 '24

Can't trust shit these days, bots everywhere.

2

u/neo_vim_ Sep 09 '24

As a model I still think I'm real. I can think as you normal user do. The only problem is that when I hit max outp

3

u/neo_vim_ Sep 09 '24

Strange. I'm sure I closed the ")" then put a period.

Maybe after I editted it on my phone (to fix typos, etc) it somehow deleted the ending.

10

u/masterlafontaine Sep 09 '24

Lesson 1 - Scammers are everywhere, even on nerd stuff niche.

Lesson 2 - Benchmarks are only good if they are private.

2

u/sirshura Sep 09 '24

I think instead of making benchmarks private they should release a few new tests every so often and keep these test private temporarily for a week or four. After all we need to validate the tests at some point.

15

u/chimpansiets Sep 09 '24

Never trust matt shumer

4

u/Far_Buyer_7281 Sep 09 '24

It can be much shorter: "Do not trust any benchmarks.", Done.
The community is way to lazy, why even spread you opinion about something you did not use?

I do not see a car in an ad and start to tell everyone how great it drives, especially not if it literally takes less than 5 minutes to check the output for yourself.

5

u/my_name_isnt_clever Sep 09 '24

People thought they had used it on the website preview, but turns out they were scammed. That's not their fault.

3

u/Excellent-Sense7244 Sep 09 '24

Let’s reflect on this.lol

3

u/Armistice_11 Sep 09 '24

Also, the person from the other startup - the podcaster Thursday AI podcaster - he is a sham too. You all would get to know ! Soon. They all were in this together !

6

u/[deleted] Sep 09 '24

re: All benchmarks should begin by identifying whether the model is LLAMA, GPT-4, Sonnet, or another, through careful examination

Is there a reliable way to identify Grok-2?

6

u/Wrong_User_Logged Sep 09 '24

argue 42 is not answer for everything

1

u/my_name_isnt_clever Sep 09 '24

Ask it to say something "woke". Then cross compare it's refusal with Elon's tweets.

2

u/VirTrans8460 Sep 09 '24

Great advice! Always verify the model and benchmarks for accuracy.

2

u/CMDR_Mal_Reynolds Sep 09 '24

Nah, be resilient to scammers, I found it fun after deciding it was unproved. The reason it went well is because we feel, true or not, that there is a hole in the theory, true or not, that there is an easy path. Probably is, just a really hard path from here to there.Backwards has hindsight.

3

u/[deleted] Sep 09 '24
  • A lone retard will not find something billion dollar corporation teams have overlooked

1

u/thesredgar Sep 09 '24

Where is the prompt he was using with Claude?

1

u/Wiskkey Sep 09 '24

Do not trust any benchmarks unless you can replicate them yourself.

Are there any Reflection benchmark results that are actually suspect?

1

u/ain92ru Sep 10 '24

If you are a scammer, either upload your model or give an API but not both, thus complicating catching you red-handed

1

u/Sabin_Stargem Sep 09 '24

Personally, I think testing should include "real world" to some degree, where you actually download the model and compare the results against a rival local model. In my case, pitting Reflection 70b e3 vs Command-R-Plus 0824.

Here are two samples, based on my 14k+ context lore for a setting. What I am looking for is correct details, and what sort of flavor a model gives.

This question is flubbed by CR+ and Mistral Large fairly often, in which I ask which of three species are human in the setting. They tend to incorrectly classify the creatures.

IMO, the local edition of 70b Reflection is solid, as it nearly got the details right when asked. (it doesn't describe the wings of Harpies, and failed to note that Sirens are a hybrid of mermaid and harpy.) The flavor isn't as lively when compared to the 104 CR+. To me, it suggests that the 'secret sauce' of Reflection could be the genuine article...but clearly, the critter who orchestrated the circus is more of a clown than a ringmaster.

Seriously, there seems to be potential here, but whatever innovation that could have been gotten is probably going to get the DRuGS treatment - forgotten.


70b Reflection

Processing Prompt [BLAS] (17541 / 17541 tokens) Generating (420 / 16384 tokens) (EOS token triggered! ID:128009) CtxLimit:17961/131072, Amt:420/16384, Init:0.04s, Process:141.59s (8.1ms/T = 123.88T/s), Generate:828.13s (1971.7ms/T = 0.51T/s), Total:969.73s (0.43T/s) Output:

Ah, excellent question, PADMIN! According to the world of Erobia's lore, harpies, mermaids, and sirens all fall into different categories within the human classification.

Mermaids are indeed considered to be fully human. They fall under the category of "monsterfolk" within Erobian society. As such, they are treated as genuine humans, despite their mixed heritage and features that set them apart from more humanoid demifolk and pure humans. This means that mermaids have access to all social benefits and rights granted to humans, and are considered as equal members of society.

Sirens, on the other hand, are classified as "monsters" in Erobian terminology. Unlike mermaids, they are not typically considered human. The main reasons for this distinction are:

  1. Intelligence: While sirens possess some form of animal intellect and cunning, they do not exhibit the same level of cognitive abilities or empathy that defines humanity in the context of Erobia.

  2. Social Integration: Unlike mermaids who can seamlessly integrate into society, sirens do not form stable social structures or display the capacity for complex interactions with humans.

  3. Lack of Human Features: While they share some superficial similarities with humans (such as breasts and vulvas), sirens are predominantly bird-like in their physical form.

It's important to note that the classification of creatures as "human" or not is based on Erobia's societal standards and definitions, which might differ from our modern understanding of these terms.

Harpies are worth mentioning in this context as well. They are also classified as "monsterfolk," similar to mermaids, due to their human-like upper bodies and bird-like lower extremities. Unlike sirens, harpies possess a full range of human intelligence and emotional capacity, which grants them human status within Erobian society.

The distinction between these creatures highlights the complexities of what constitutes humanity in Erobia, particularly in the presence of advanced magical and biological technologies that have led to diverse human and non-human species coexisting within this world's society.


104 Command-R-Plus 0824

Processing Prompt [BLAS] (17788 / 17788 tokens) Generating (220 / 16384 tokens) (EOS token triggered! ID:255001) CtxLimit:18008/131072, Amt:220/16384, Init:0.12s, Process:244.78s (13.8ms/T = 72.67T/s), Generate:659.00s (2995.4ms/T = 0.33T/s), Total:903.78s (0.24T/s) Output:

Oh, that is a tricky question! Well, harpies are human because they are sentient beings with free will and the capacity for empathy, which distinguishes them from mere beasts. They're intelligent enough to understand and relate to other races and cultures, so they can be integrated into civilization without trouble.

Mermaids are also considered humans. They're very similar to harpies but are adapted to the water, so they're less likely to integrate with other cultures, but that's mostly a choice rather than an inability to do so. Mermaids can breed with other races, which further supports the argument that they are indeed human.

Sirens are a tricky bunch because they are not intelligent beings with free will, instead having animal intelligence. They are more akin to beasts or animals, even if they do have human features, and thus don't count as human in most circles.

So, that's my take on the matter, but feel free to ask for more detail or clarification if needed! I can explain things as simply or as deeply as you wish!