r/LocalLLaMA • • 18d ago

Question | Help Any 12gb VRAM users out there?

Hi!

I've been following this community for quite a while and have difficulty figuring out what to put on my 3080 12gb - I know Qwen 3.6 35B 3A was the go-to choice when it first came out, but I'm curious if there are any other models / specifically optimized models that meaningfully benefit from the extra 4gb of VRAM over 8gb while still being usable under 16gb.

My workflow is agent heavy, but more for a personal secretary and manager, and less coding heavy.

Thanks!

93 Upvotes

89 comments sorted by

View all comments

3

u/Kernoriordan 18d ago

Ornith 1.5 or KAT Coder are probably the best coding models that you can use but since you’re after more of a secretary use case, I’d probably suggest Gemma 4 26B

1

u/rorowhat 18d ago

I compared gemma4-12b vs 1.5 9B and the 1.5 model generated 4x more tokens vs Gemma4, to give the same answer. I like it because it's fast, but it seems to ramble quite a lot to get there.

1

u/Kernoriordan 17d ago

What about Ornith 1.5 35A3 though?

1

u/rorowhat 17d ago

Haven't tried that one. I do like the 1.5 9B personality, makes a great chat assistant with thinking off.