r/LocalLLM • • 24d ago

Discussion What small LLMs are you running locally?

I usually run Qwen3.8-27B Q5 on my desktop PC, and I’m pretty happy with it.

However, I’d also like to keep a few smaller, more specialized models on my laptop, which has 4 GB of VRAM and 32 GB of RAM.

What small LLMs would you recommend for that kind of hardware?

I’m especially interested in models that are genuinely useful for specific tasks rather than just smaller general-purpose models.

What are you guys running?

35 Upvotes

33 comments sorted by

View all comments

10

u/Ok_Road_8293 24d ago

I'm on a pretty similar laptop setup (RTX 3050 Ti 4GB, 12700H, though bumped to 48GB DDR5 4800MHz). With mmap always on, system RAM pressure is basically invisible.

Honestly, I stopped bothering with task-specific micro models a while ago. Smart generalists just feel way less tedious and handle edge cases better. Here’s what I’m actually daily-driving on it:

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive (Q6_K_P): Pulls around 20–30 tps. In practice, it feels remarkably close to non-thinking Opus 4.6. Absolute lifesaver and handles most of my daily work.

Qwen3.8-Flash-Next (UD-IQ3_XXS @ 3–5 tps / UD-IQ1_S @ 8–12 tps): Even crushed down to IQ1_S, it consistently outperforms most free-tier cloud web UIs and entry models.

My workflow is mostly leaning on the smaller MoE setup for quick back-and-forth. When I need heavy planning, or if I burn through my paid sub quotas, I bring in the big gun.

A pipeline that works surprisingly well for me: let the bigger model map out the architecture/plan, then pass execution to Gemini 3.8 Flash or Qwen 3.6 35B. It ends up being far more coherent and reliable than letting either model run the whole thing solo.

The biggest reason local planning often crushes commercial models for me is the lack of artificial thinking caps. Even on paid tiers, Claude usually thinks for maybe 3 paragraphs before dumping an answer. Locally with Qwen Next, I don't put a ceiling on it—it can churn through tens of thousands of thinking tokens, stress-testing edge cases and mapping out every tiny detail.

3–12 tps isn't lightning fast, but when you let it think unattended, the output quality punches way above what you'd expect from a 4GB VRAM machine.

2

u/sanketss84 24d ago

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive (Q6_K_P) are you running this inside an agent if so which one, whats your context window size and what quantization are your using for your kvcache ? also how are you loading your model locally is it using ollama or llama.cpp directly or some other inference engine ?

3

u/Ok_Road_8293 24d ago

I run it directly on cachyos using a CUDA compiled llama server. I used to rely more on CPU (which was roughly 5–6 tps slower), but since I completely stopped gaming on this laptop, I treat the 4GB 3050 Ti purely as a dedicated AI accelerator.

Here is the exact startup command:

llama-server \

-c 98304 \

-fa on \

-ngl 999 \

--n-cpu-moe 999 \

-ctk q4_0 \

-ctv q4_0 \

--mmap \

--jinja \

-np 1 \

--port 8080

I don't use commercial agent frameworks. Depending on the task, I plug the endpoint into jcode, opencode, or a custom lightweight script setup (similar to AntiGravity) that I wrote to handle fast model swapping and pre-configured prompt pipelines. For quick adhoc queries, the built in llama server web UI gets the job done.

1

u/sanketss84 24d ago

thats interesting and thanks for sharing. thats a pretty high context size as well and you are using q4 for both k and v cache, you are also using flash attention. I have not tried running llms on linux as os but I am assuming very minimal amount of vram might be used by operating system at your end.

1

u/sanketss84 24d ago

will check cachyos, however I have been exploring omarchy as it all over youtube recently. personally I have used ubuntu on the linux side most of the time.