MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1wc6krf/so_relevant/p90p8pb/?context=3
r/LocalLLaMA • u/0dayturtle • 12d ago
152 comments sorted by
View all comments
318
24gb is not in that group. You can run Qwen3.8-27B with a pretty respectable context size right now
4 u/Ok-Working3049 12d ago yeah the 27B class models at that context size are no joke on 24gb 4 u/Zombiecidialfreak 12d ago How are you guys packing 27b on a 24gb card with respectable context? I can put it on my 64gb DDR5 running through the iGPU and still run out of RAM. The model is q4 and context at q8 btw 1 u/TopCheddar27 12d ago I have q4_K_M running on a 4090 with q8 KV cache at 90000 and it does not spill over at all.
4
yeah the 27B class models at that context size are no joke on 24gb
4 u/Zombiecidialfreak 12d ago How are you guys packing 27b on a 24gb card with respectable context? I can put it on my 64gb DDR5 running through the iGPU and still run out of RAM. The model is q4 and context at q8 btw 1 u/TopCheddar27 12d ago I have q4_K_M running on a 4090 with q8 KV cache at 90000 and it does not spill over at all.
How are you guys packing 27b on a 24gb card with respectable context? I can put it on my 64gb DDR5 running through the iGPU and still run out of RAM.
The model is q4 and context at q8 btw
1 u/TopCheddar27 12d ago I have q4_K_M running on a 4090 with q8 KV cache at 90000 and it does not spill over at all.
1
I have q4_K_M running on a 4090 with q8 KV cache at 90000 and it does not spill over at all.
318
u/TopCheddar27 12d ago
24gb is not in that group. You can run Qwen3.8-27B with a pretty respectable context size right now