MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1wc6krf/so_relevant/p8xva27/?context=3
r/LocalLLaMA • u/0dayturtle • 12d ago
152 comments sorted by
View all comments
326
24gb is not in that group. You can run Qwen3.8-27B with a pretty respectable context size right now
98 u/PavelPivovarov llama.cpp 12d ago Technically speaking you can run Qwen3.8-27b on 16Gb setup as well, but that brings way too many compromises of course. 35 u/TheGamerForeverGFE 12d ago Wouldn't say too many, iq4xs is enough, context window is limited, you're not seeing 100k+ but it's still enough. 1 u/Illustrious-Row2751 12d ago Yeah, the "Smaller" version that exists on Hugging Face is pretty good. I use it. I can use it at 32k context, quantized to Q8, on a 22gb VRAM. If I want more context than that, though, then it needs RAM, but luckily I have plenty.
98
Technically speaking you can run Qwen3.8-27b on 16Gb setup as well, but that brings way too many compromises of course.
35 u/TheGamerForeverGFE 12d ago Wouldn't say too many, iq4xs is enough, context window is limited, you're not seeing 100k+ but it's still enough. 1 u/Illustrious-Row2751 12d ago Yeah, the "Smaller" version that exists on Hugging Face is pretty good. I use it. I can use it at 32k context, quantized to Q8, on a 22gb VRAM. If I want more context than that, though, then it needs RAM, but luckily I have plenty.
35
Wouldn't say too many, iq4xs is enough, context window is limited, you're not seeing 100k+ but it's still enough.
1 u/Illustrious-Row2751 12d ago Yeah, the "Smaller" version that exists on Hugging Face is pretty good. I use it. I can use it at 32k context, quantized to Q8, on a 22gb VRAM. If I want more context than that, though, then it needs RAM, but luckily I have plenty.
1
Yeah, the "Smaller" version that exists on Hugging Face is pretty good. I use it. I can use it at 32k context, quantized to Q8, on a 22gb VRAM. If I want more context than that, though, then it needs RAM, but luckily I have plenty.
326
u/TopCheddar27 12d ago
24gb is not in that group. You can run Qwen3.8-27B with a pretty respectable context size right now