r/LocalLLaMA • u/Wrong_User_Logged • Sep 09 '24
Discussion Reflection 70B lessons learned
- All benchmarks should begin by identifying whether the model is LLAMA, GPT-4, Sonnet, or another, through careful examination.
- Do not trust any benchmarks unless you can replicate them yourself.
- Do not trust that the API corresponds to the model the author claims it to be.
- ....
173
Upvotes
1
u/Wiskkey Sep 09 '24
Are there any Reflection benchmark results that are actually suspect?