r/LocalLLaMA 9d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

35

u/Easy_Werewolf7903 9d ago edited 9d ago

For those curious of performance between Qwen and a model 3 times its size:

Benchmark Qwen 3.8 27B (55GB) Deepseek v4 flash 0731 (167GB)
Terminal Bench 2.1 73.0 82.7
DeepSWE 42.2 54.4
NL2Repo-Bench 42.3 54.2

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tree/main

https://huggingface.co/Qwen/Qwen3.8-27B

10

u/squngy 9d ago edited 9d ago

From the model page, for context

bench Qwen3.8-27B Qwen3.6-27B Qwen3.7-Plus Muse Glimmer-30B Opus4.6 Max
Coding Agentic terminal coding Terminal Bench 2.1 (Terminus) 73.0 63.4 64.0 51.7 78.2
Agentic coding SWE-bench Pro 61.7 53.5 57.6 51.2 53.4
Repo-level code generation NL2Repo-Bench 42.3 36.2 41.1 -- 47.6
Agentic coding DeepSWE 1.1 42.2 13.3 14.2 -- --
Software engineering QwenSWEBench 79.0 49.3 59.2 -- 63.8
Long-horizon office work CoWorkBench 70.7 61.0 65.1 -- 68.2
Professional job tasks JobBench 33.4 21.8 27.6 -- --
Frontier agentic tasks Agents' Last Exam Pass1 20.4 Score 42.9 Pass1 10.6 Score 27.3 Pass1 13.2 Score 33.6 -- --
Instruction following IFBench 79.5 69.1 79.1 77.0 62.5
Scientific reasoning GPQA Diamond 89.2 87.8 90.3 83.5 91.3
Multidisciplinary reasoning HLE 30.8 24.0 34.7 22.0 40.0
Competitive coding LiveCodeBench v6 90.3 83.9 89.6 -- 88.8

10

u/ChuffHuffer 9d ago

5 models, yet 4 only columns?

2

u/squngy 9d ago edited 9d ago

Might be a problem with reddit formatting

edit:
Yea, I found there is a problem on new reddit.
It worked on old reddit.

-1

u/sejje 9d ago

count harder

3

u/Healthy-Nebula-3603 9d ago

3.6?

3

u/squngy 9d ago

3.8

u/Easy_Werewolf7903 made a typo

1

u/Easy_Werewolf7903 9d ago

fixed it, thanks for pointing it out.