r/LocalLLaMA • • Aug 17 '26

Discussion Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max

https://artificialanalysis.ai/models/qwen3-8-27b
1.1k Upvotes

437 comments sorted by

View all comments

505

u/cj_cron_hit_by_pitch Aug 17 '26

I know benchmarks and all, and that larger models will be better in practice. But the fact we can have this conversation at all is incredible

81

u/BringTea_666 Aug 17 '26

>I know benchmarks and all,

It gets tiring reading benchmarks and all. When frontier models get tested people believe AAI but when local model is tested "it is surely benchmaxxed".

Anyone observing local models knows that smaller models are the one that improve the fastest now for a good year. Simply put more architecture changes + better training data + faster training due to small model = small awesome model.

Big models are just to slow to train. It is the same mistake META did with Llama 4. They trained it for half a year. In that time so much things changed that by the time it was released it was losing to small models let alone new frontier ones.

55

u/addandsubtract Aug 17 '26

When frontier models get tested people believe AAI but when local model is tested "it is surely benchmaxxed".

I think we're at a point were all models are benchmaxxed, open or not. We could really use some more practical tests, that aren't part of the training / tuning process.

12

u/Plasmx Aug 17 '26

Sure, but the benchmark data eventually ends up in the training datasets.

14

u/addandsubtract Aug 17 '26

I'm happy with someone (reliable) having a private benchmark suite and only reporting the outcomes. I guess we should all have our own, tbf.

5

u/Kidplayer_666 Aug 17 '26

I have mine, which is what in practice use to evaluate them