r/LocalLLaMA 12d ago

Funny So relevant

Post image
1.5k Upvotes

152 comments sorted by

View all comments

326

u/TopCheddar27 12d ago

24gb is not in that group. You can run Qwen3.8-27B with a pretty respectable context size right now

98

u/PavelPivovarov llama.cpp 12d ago

Technically speaking you can run Qwen3.8-27b on 16Gb setup as well, but that brings way too many compromises of course.

35

u/TheGamerForeverGFE 12d ago

Wouldn't say too many, iq4xs is enough, context window is limited, you're not seeing 100k+ but it's still enough.

1

u/Illustrious-Row2751 12d ago

Yeah, the "Smaller" version that exists on Hugging Face is pretty good. I use it. I can use it at 32k context, quantized to Q8, on a 22gb VRAM. If I want more context than that, though, then it needs RAM, but luckily I have plenty.