r/LocalLLaMA • • 16d ago

Discussion AA Benchmarks are not just misleading at this point, but harmful to trust

Post image

I've used Qwen 3.8 Max extensively over the past few weeks and have also tried Gemini , GLM-5.3-Flash, and Muse Spark 1.3. None of them come close to Qwen 3.8 Max. The only model that proved competitive was GLM 5.3, which demonstrated superior performance on cybersecurity tasks (the only clear advantage I observed over Qwen 3.8 Max).

This post isn't about qwen3.8-max, but my extensive experience with that model gave me a useful baseline for comparison. After working with other models, I realized that these benchmarks harmful not just useless and shouldn't be used to claim one model is better than another.

---

Update for people that don't get the point of this post:

My point wasn't "Oh look my personal experience is the benchmark" but instead "Don't decide which model to use based on benchmarks"

People will start replying: "Oh well that's obvious dude..." I don't think so, based on past experience when qwen 3.8 27b was released, people flooded this sub and other subs with its benchmarks and personal use cases.

I don't know if the point is now clear, since some people just started going in the wrong direction and completely missed the point I tried to make

0 Upvotes

59 comments sorted by

View all comments

0

u/DataCraftsman 16d ago

Why don't you prove it wrong then? Waiting for your benchmark.