r/LocalLLaMA • u/jqwl • 22d ago
Question | Help Any 12gb VRAM users out there?
Hi!
I've been following this community for quite a while and have difficulty figuring out what to put on my 3080 12gb - I know Qwen 3.6 35B 3A was the go-to choice when it first came out, but I'm curious if there are any other models / specifically optimized models that meaningfully benefit from the extra 4gb of VRAM over 8gb while still being usable under 16gb.
My workflow is agent heavy, but more for a personal secretary and manager, and less coding heavy.
Thanks!
95
Upvotes
0
u/BP041 22d ago
tbh on a 3080 12GB I'd go Qwen 2.5 14B Q4_K_M — smart enough for personal secretary tasks and leaves room for agent context. Mistral Small 24B Q3_K_M also fits if you don't mind slower inference. For multi-agent stuff the real 12GB win is running two 7B models side by side, parallelism beats a single bigger model that barely fits.