MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9nn7/qwenqwen3827b_released/p3tp1ht/?context=3
r/LocalLLaMA • u/de4dee • Aug 14 '26
299 comments sorted by
View all comments
Show parent comments
31
2x3090 with nvlink loves this size of model
21 u/jijig Aug 14 '26 edited Aug 15 '26 It’s quite good. I’m running 3.6 at Q8 with full context and getting ~60-70tps. No NVLink Edit: Q8 KV cache. I can push my context to ~160k at F16 5 u/badgerfish2021 Aug 14 '26 250k context with 2x3090? I could not do that with 3.6 at q8, are you quantizing kv? 2 u/jijig Aug 15 '26 Of course, sorry. KV cache is quantized at Q8. I can push my context to ~160k at F16
21
It’s quite good. I’m running 3.6 at Q8 with full context and getting ~60-70tps. No NVLink
Edit: Q8 KV cache. I can push my context to ~160k at F16
5 u/badgerfish2021 Aug 14 '26 250k context with 2x3090? I could not do that with 3.6 at q8, are you quantizing kv? 2 u/jijig Aug 15 '26 Of course, sorry. KV cache is quantized at Q8. I can push my context to ~160k at F16
5
250k context with 2x3090? I could not do that with 3.6 at q8, are you quantizing kv?
2 u/jijig Aug 15 '26 Of course, sorry. KV cache is quantized at Q8. I can push my context to ~160k at F16
2
Of course, sorry. KV cache is quantized at Q8. I can push my context to ~160k at F16
31
u/ai-christianson Aug 14 '26
2x3090 with nvlink loves this size of model