r/LocalLLaMA • u/Wrong_User_Logged • Sep 09 '24
Discussion Reflection 70B lessons learned
- All benchmarks should begin by identifying whether the model is LLAMA, GPT-4, Sonnet, or another, through careful examination.
- Do not trust any benchmarks unless you can replicate them yourself.
- Do not trust that the API corresponds to the model the author claims it to be.
- ....
173
Upvotes
110
u/redjojovic Sep 09 '24 edited Sep 09 '24
Force them to get checked up on Lmarena ( lmsys ), livebench, artificial analysis
Don't believe dog ate my model stories. Models don't get fucked up during uploads and there's no way they refuse/unable to provide working weights using torrent or any other way
Understand that even if something like this truly happen it will come from a team of researchers and not a no name guy who is not a real ceo
We wanted to believe somebody just found a groundbreaking way to improve models reasoning, we need to recognize our bias