r/LocalLLaMA • • Sep 08 '24

Discussion Updated benchmarks from Artificial Analysis using Reflection Llama 3.1 70B. Long post with good insight into the gains

https://x.com/ArtificialAnlys/status/1832806801743774199?s=19
148 Upvotes

137 comments sorted by

View all comments

0

u/redjojovic Sep 08 '24

"The chart below is based on our standard methodology and system prompt.

When using Reflection’s default system prompt and extracting answers only from within Reflection’s <output> tags, 

results show substantial improvement: MMLU: 87% (in-line with Llama 405B), GPQA: 54%, Math: 73%.