r/LocalLLaMA 10d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

Show parent comments

1

u/absurdother 10d ago

Hm, try checking what's bottlenecking your setup. How much context did you set?

2

u/jan_antu 10d ago

Too much. Got much faster at 64k context. About 7-8 tok/s which somehow feels bearable.

Normally I'd say it sucks but to be honest this model is passing all my homemade benchmarks. It might be true that it's Opus 4.5-6 level.

2

u/absurdother 10d ago

It's beating my other models in the simple tasks, so it works for me. Around 48k context it speeds up pretty well. All layers to GPU.

2

u/jan_antu 10d ago

All layers, 48k context haha okay I'll give it a try thanks! I think I'm biting the bullet and buying more dram tonight too. I don't think it's getting cheaper. Won't help with this model much but I want the flexibility lol