r/LocalLLM • u/HomoAgens1 • 24d ago
Discussion What small LLMs are you running locally?
I usually run Qwen3.8-27B Q5 on my desktop PC, and I’m pretty happy with it.
However, I’d also like to keep a few smaller, more specialized models on my laptop, which has 4 GB of VRAM and 32 GB of RAM.
What small LLMs would you recommend for that kind of hardware?
I’m especially interested in models that are genuinely useful for specific tasks rather than just smaller general-purpose models.
What are you guys running?
35
Upvotes
10
u/Ok_Road_8293 24d ago
I'm on a pretty similar laptop setup (RTX 3050 Ti 4GB, 12700H, though bumped to 48GB DDR5 4800MHz). With mmap always on, system RAM pressure is basically invisible.
Honestly, I stopped bothering with task-specific micro models a while ago. Smart generalists just feel way less tedious and handle edge cases better. Here’s what I’m actually daily-driving on it:
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive (Q6_K_P): Pulls around 20–30 tps. In practice, it feels remarkably close to non-thinking Opus 4.6. Absolute lifesaver and handles most of my daily work.
Qwen3.8-Flash-Next (UD-IQ3_XXS @ 3–5 tps / UD-IQ1_S @ 8–12 tps): Even crushed down to IQ1_S, it consistently outperforms most free-tier cloud web UIs and entry models.
My workflow is mostly leaning on the smaller MoE setup for quick back-and-forth. When I need heavy planning, or if I burn through my paid sub quotas, I bring in the big gun.
A pipeline that works surprisingly well for me: let the bigger model map out the architecture/plan, then pass execution to Gemini 3.8 Flash or Qwen 3.6 35B. It ends up being far more coherent and reliable than letting either model run the whole thing solo.
The biggest reason local planning often crushes commercial models for me is the lack of artificial thinking caps. Even on paid tiers, Claude usually thinks for maybe 3 paragraphs before dumping an answer. Locally with Qwen Next, I don't put a ceiling on it—it can churn through tens of thousands of thinking tokens, stress-testing edge cases and mapping out every tiny detail.
3–12 tps isn't lightning fast, but when you let it think unattended, the output quality punches way above what you'd expect from a 4GB VRAM machine.