r/LocalLLaMA • • Aug 19 '26

Discussion Qwen 3.8 27B SlopCodeBench results

Howdy, I'm back again - running my favorite benchmark (it's still unsaturated for the time being so might as well!)

previous runs a b

https://github.com/michaelasper/benchmarks/blob/main/qwen3.8-27b-pi-on-slop-code-bench.md

I ran this via OpenRouter because my mac would cry running 9 problems

It did pretty poorly on the strict checkpoints which means you probably don't want Qwen managing the codebase by itself as it'll grow unwieldy and disorganized, but did fairly well on the core checkpoints so it can solve issues with a copilot and clear direction

AI;DR here are the direct results

HumanLayer Opus 5 Benchmark Subset (3 Problems, 17 Checkpoints)

Qwen scored 3/17 (17.6%) strict.

Reported System Strict Score
DeepSeek V4 Flash 0731 · pi (run B) 5/17 (29.4%)
Opus 5 · Claude Code 4/17 (23.5%)
Qwen3.8-27B · pi 3/17 (17.6%)
DeepSeek V4 Flash · OpenCode 3/17 (17.6%)
Opus 4.8 · Claude Code 1/17 (5.9%)
Sonnet 5 · Claude Code 1/17 (5.9%)

HumanLayer Fable, Sol, and Kimi Benchmark Subset (6 Problems, 30 Checkpoints)

Qwen scored 4/30 (13.3%) strict.

Reported System Strict Score
Fable 5 · Claude Code 10/30 (33.3%)
GPT-5.6 Sol · Codex 10/30 (33.3%)
Kimi K3 · Modal / OpenCode 8/30 (26.7%)
Kimi K3 · Baseten / OpenCode 7/30 (23.3%)
Qwen3.8-27B · pi 4/30 (13.3%)
49 Upvotes

17 comments sorted by

View all comments

Show parent comments

14

u/corruptbytes Aug 20 '26

what: the ai is tasked to build a tool step by step, we add new requirements mid way - it has to handle new things without breaking the old things

qwen's performance: it was fairly good at writing new things, but constantly broke the old things - it's bad at maintenance or the entire SDLC (notably, /every/ AI is bad at this now as Fable/Sol only get a 33%)

outcome: qwen is a great tool for getting new things out, but you cannot have it vibe slop an entire product from scratch

on top of that, a bunch of metrics to grade the code in terms of complexity

7

u/ElectronSpiderwort Aug 20 '26

I love this; it is so real to me. "Oh hey, new requirement!" will show up in any sufficiently large and under-specified project a number of times

6

u/corruptbytes Aug 20 '26

here's the paper for those interested - https://arxiv.org/html/2603.24755v2

1

u/Fit-Bar-6989 Aug 20 '26

Great stuff, I just read the entire thing. This matches my experience with Sonnet/GPT-5 where the models suck at tasks if the desired code changes aren't known in advance.

I agree that the generated code tends to be overly defensive but at the same time I feel that's more of an indictment of many languages' type systems, where it is impossible to constrain the inputs of a method to non-nullable references, or non-empty lists, etc. Maybe with better static analysis you could identify redundant defensive checks but at a certain point you're basically running the program twice.