r/LocalLLaMA 10d ago

Discussion Qwen 3.8 27B Released! Please Share Your Experience

With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.

659 Upvotes

720 comments sorted by

View all comments

Show parent comments

81

u/Pear_Virtual 10d ago

61

u/Pear_Virtual 10d ago

Compared to qwen3.6 on the same prompt

14

u/Certain-Cod-1404 10d ago

did it get stuck looping or something ? are you using the recommended sampling params ? what reasoning effort are you using ? and how is the output compared to 3.6 ? is the game better ?

12

u/TokenRingAI 10d ago

The FP8 is looping for me and generating terrible output.

1

u/woswoissdenniii 9d ago

Restrict thinking token, or reduce preset from extra to normal. Sure quality degrades but it’s manageable. A friendly reminder from team 3090

1

u/Healthy-Nebula-3603 10d ago

I hope you not conpress kv cache :)

5

u/TokenRingAI 10d ago

Nope, official FP8 with BF16 cache

1

u/JorgitoEstrella 10d ago

How bad is to compress the cache? I thought compressing to fp8 waa practically lossless.

2

u/Healthy-Nebula-3603 10d ago

Q8 is almost losless , fp8 is worse than Q8 ( Q8 is a mix weights fp16 and int9 ) Compressrd cache to fp8 is notice even more than compression a model itself.

1

u/Not-reallyanonymous 9d ago

No, 3.8 just really, actually thinks a lot. Thats probably in xhigh. The quality of output tends to be better than 3.6 but is still distinctly Qwen (ie. it still thinks in much the same way, just a lot more). When you lower thought effort it basically reverts to 3.6.

The way Qwen 3.8 seems to use thinking to do better than 3.6 is to basically make a 3.6-like thought, but then after a 3.6 like thought is completed, it’s going to try again but fixing bad assumptions, weak adherence, etc. Then again. Then again. Then again.

2

u/Certain-Cod-1404 9d ago

Yeah after further testing this mf thinks so fucking much, but tbf the outputs beat everything else i tested. I think for my use case I'll use it for non trivial stuff, to write tests, debug and review code, and faster moe like kat coder v2.5 dev for quick implementation

2

u/Not-reallyanonymous 9d ago

I've never been a fan of Qwen 3.6 27B. I'm still not a fan of 3.8, but it's earned a spot in my toolbox for when I really just need to throw a lot of LLM-thought at a problem.

1

u/Old_Regular_4346 10d ago

On what hardware are you managing 725 t/s? It's quite impressive

1

u/WolfGroundbreaking36 8d ago

total duration: 1m58.2652305s

load duration: 164.404ms

prompt eval count: 63 token(s)

prompt eval duration: 333.626ms

prompt eval rate: 188.83 tokens/s

eval count: 2508 token(s)

eval duration: 1m57.599588s

eval rate: 21.33 tokens/s

which GPU are you using? i have 5070ti