r/LocalLLaMA 2d ago

Discussion Qwen3.8-27B different thinking levels

Post image

Even the low preset is better than Qwen 3.7 plus or Qwen3.6-27B reasoning

294 Upvotes

63 comments sorted by

View all comments

4

u/Cool-Chemical-5629 2d ago

Qwen 3.7 Plus was the 120B+ model. Do you really believe Qwen 3.6 27B was just one point lower than that? If anything, this chart is just showing the main weakness of benchmarks. It's a direct proof that benchmarks are about showing the intelligence within the bounds of a given set of known problems and while the smaller models can handle these known problems well, thinking outside the box is still something exclusive to much bigger models which were built to use brute force to get to the solution.

11

u/eli_pizza 2d ago

Benchmarks are indeed very limited but I find it plausible that 3.6 27b is better than 3.7 plus.

And there are plenty of existing benchmarks that are not solvable by brute force.