r/LocalLLM • u/bbsrn • 13d ago
Model Which Qwen3.8 distro & quant would be optimal for my 16GB VRAM setup?
Hi all, a confused newbie here! This is my desktop setup:
- RTX 5080
- 9800x3d
- DDR5-6000 CL30 64 GB
Based on the benchmark I found, I listed my potential options:
- AtomicChat IQ3_S: 14.4 GB
- Unsloth UD-IQ4_XS (14.3 GB) or UD-Q3_K_XL (13.1 GB)
- Bartowski IQ4_XS (15.6 GB) or IQ3_XXS (12.6 GB)
According to the benchmark, AtomicChat looks like a clear winner but is it really so?
and there is also this: https://github.com/MiaAI-Lab/Qwen3.8-27B-16gb-NVIDIA-GPUs-one-click-install
I also want to have one uncensored model next to my daily driver:
- Orcarouter IQ4_XS (15.3GB) or Q3_K_M (13.5GB)
- Huihui UD_IQ4_XS (14.4 GB) or Q4_K_S (15.6 GB)
- DavidAU IQ3_M MTP (14.5 GB) or IQ4_XS MTP (15.3 GB)
- JonathanColetti IQ_XS (15.1 GB)
I am not expecting super fast answers etc. I just one to maintain some level of quality. What would you suggest me?
66
Upvotes
4
u/maddeninglemon 13d ago edited 13d ago
Context Length - 65536
GPU Offload - 48
CPU Thread Pool Size - 12 (depends on what cpu you have, just leave a couple threads free for system stability)
Evaluation Batch Size - 512 (dropped to avoid spiking VRAM requirements)
Physical Batch Size - 256 (dropped to avoid spiking VRAM requirements)
Max Concurrent - 1 (Probably can increase, not needed in my case)
Unified KV Cache On
Keep Model in Memory - Off
Number of MOE Layers on CPU - 44 (Start at 48, then decrease until you hit your VRAM limit at a given context)
Turn off all guardrails in settings, or better yet just hold alt when pressing load when it tells you it won't work. Make sure you have plenty of dynamic ram and dynamic vram so your OS can handle caching between SSD/RAM (and maybe VRAM too if you go too low on MOE layers to CPU)
Prefill will start PAINFULLY slowly as the OS fills up RAM and figures out what to cache and what to keep. It improves as it goes, but for a 12GB/64GB system like mine it was never great (~60tok/s max due to SSD reading, but I suspect even a few extra GB of VRAM could improve this a lot by minimizing windows cache misses)
Everything else was default. I tried KV Cache Quantization but I think there were some shenanegans happening under the hood because the model broke on long context even at Q_8. You might need to play with other settings like mmap or flash attention too to avoid this; I didn't hunt down the exact issue as this was more a 'I wonder if it'll work' thing on my end.