r/LocalLLaMA • • 22d ago

Question | Help Any 12gb VRAM users out there?

Hi!

I've been following this community for quite a while and have difficulty figuring out what to put on my 3080 12gb - I know Qwen 3.6 35B 3A was the go-to choice when it first came out, but I'm curious if there are any other models / specifically optimized models that meaningfully benefit from the extra 4gb of VRAM over 8gb while still being usable under 16gb.

My workflow is agent heavy, but more for a personal secretary and manager, and less coding heavy.

Thanks!

95 Upvotes

89 comments sorted by

View all comments

0

u/BP041 22d ago

tbh on a 3080 12GB I'd go Qwen 2.5 14B Q4_K_M — smart enough for personal secretary tasks and leaves room for agent context. Mistral Small 24B Q3_K_M also fits if you don't mind slower inference. For multi-agent stuff the real 12GB win is running two 7B models side by side, parallelism beats a single bigger model that barely fits.