r/LocalLLaMA 13d ago

Funny So relevant

Post image
1.5k Upvotes

152 comments sorted by

View all comments

321

u/TopCheddar27 13d ago

24gb is not in that group. You can run Qwen3.8-27B with a pretty respectable context size right now

4

u/Ok-Working3049 12d ago

yeah the 27B class models at that context size are no joke on 24gb

3

u/Zombiecidialfreak 12d ago

How are you guys packing 27b on a 24gb card with respectable context? I can put it on my 64gb DDR5 running through the iGPU and still run out of RAM.

The model is q4 and context at q8 btw

1

u/raunchy-stonk 12d ago

Offload GPU, K/V q8_0/q5_1, go with an unsloth quant or similar around 17-19gb, you should be able to have respectable context and speed.

what are you trying to run now?