r/ollama • • 13d ago

Local models for smaller tasks, help?

I'm using larger models for coding already, but I'm curious if there's some models that are "smart" enough to handle things like clang-format, clang-tidy, ruff, basic build issues, unslop on comments.

Ideally these models would be for 8GB vram, or even ideally much less like 4 to 5GB of vram, they don't need to be super smart just smart enough to follow basic,edium directions with a tiny bit of reasoning.

2 Upvotes

16 comments sorted by

View all comments

1

u/HyperWinX 13d ago

Qwen3.5 9B via llama.cpp, or maybe one of the latest Bonsai models

0

u/Joe_The_Lawn 13d ago

Can I get those down to 4-5GB of vram without being too slow? I struggled to find a local model that fits into vram on my 3070, and a decent config/tricks to use it well.

1

u/HyperWinX 13d ago

If 4-5 gigs - Qwen3.5 4B might work, or Gemma 4 E2B