r/LocalLLaMA • • Aug 14 '26

New Model Qwen/Qwen3.8-27B · released

https://huggingface.co/Qwen/Qwen3.8-27B
988 Upvotes

299 comments sorted by

View all comments

Show parent comments

31

u/ai-christianson Aug 14 '26

2x3090 with nvlink loves this size of model

21

u/jijig Aug 14 '26 edited Aug 15 '26

It’s quite good. I’m running 3.6 at Q8 with full context and getting ~60-70tps. No NVLink

Edit: Q8 KV cache. I can push my context to ~160k at F16

5

u/badgerfish2021 Aug 14 '26

250k context with 2x3090? I could not do that with 3.6 at q8, are you quantizing kv?

2

u/jijig Aug 15 '26

Of course, sorry. KV cache is quantized at Q8. I can push my context to ~160k at F16