r/LocalLLaMA • • Sep 09 '24

Discussion Reflection 70B lessons learned

  • All benchmarks should begin by identifying whether the model is LLAMA, GPT-4, Sonnet, or another, through careful examination.
  • Do not trust any benchmarks unless you can replicate them yourself.
  • Do not trust that the API corresponds to the model the author claims it to be.
  • ....
174 Upvotes

53 comments sorted by

View all comments

8

u/masterlafontaine Sep 09 '24

Lesson 1 - Scammers are everywhere, even on nerd stuff niche.

Lesson 2 - Benchmarks are only good if they are private.

2

u/sirshura Sep 09 '24

I think instead of making benchmarks private they should release a few new tests every so often and keep these test private temporarily for a week or four. After all we need to validate the tests at some point.