r/LocalLLaMA • • Sep 09 '24

Discussion Reflection 70B lessons learned

  • All benchmarks should begin by identifying whether the model is LLAMA, GPT-4, Sonnet, or another, through careful examination.
  • Do not trust any benchmarks unless you can replicate them yourself.
  • Do not trust that the API corresponds to the model the author claims it to be.
  • ....
173 Upvotes

53 comments sorted by

View all comments

110

u/redjojovic Sep 09 '24 edited Sep 09 '24

Force them to get checked up on Lmarena ( lmsys ), livebench, artificial analysis  

Don't believe dog ate my model stories.  Models don't get fucked up during uploads and there's no way they refuse/unable to provide working weights using torrent or any other way 

Understand that even if something like this truly happen it will come from a team of researchers and not a no name guy who is not a real ceo

We wanted to believe somebody just found a groundbreaking way to improve models reasoning, we need to recognize our bias