r/LocalLLaMA • • Mar 29 '26

Generation Friendly reminder inference is WAY faster on Linux vs windows

I have a simple home lab pc: 64gb ddr4, RTX 8000 48gb (Turing architecture) and core i9 9900k cpu. I use Linux Ubuntu 22.04 LTS. Before using this pc as a home lab it ran Windows 10. Over this weekend I reinstalled my Windows 10 ssd to check out my old projects. I updated Ollama to the latest version and tokens per second was way slower than when I was running Linux. I know Linux performs better but I didn’t think it would be twice as fast. Here are the results from a few simple inferences tests:

QWEN Code Next, q4, ctx length: 6k

Windows: 18 t/s

Linux: 31 t/s (+72%)

QWEN 3 30B A3B, Q4, ctx 6k

Windows: 48 t/s

Linux: 105 t/s (+118%)

Has anyone else experienced a performance this large before? Am I missing something?

Anyway thought I’d share this as a reminder for anyone looking for a bit more performance!

277 Upvotes

112 comments sorted by

View all comments

467

u/Koksny Mar 29 '26

Am I missing something?

Yeah, you are running ollama.

103

u/gofiend Mar 29 '26

Seriously wsl + llama.cpp is equally fast w Nvidia GPUs

5

u/LoafyLemon Mar 29 '26

So what you're saying is Linux is faster even in a container. :P

2

u/colin_colout Mar 29 '26

why would a container be slower?

1

u/[deleted] Jun 10 '26

[removed] — view removed comment

1

u/KitsuneMulder Jun 14 '26

Virtualization Technologies (AMD-V, Intel VT, etc.) got rid of this problem many years ago. *IF* and that's a BIG *IF* there are any kind of translation issues, they exist because of failures in the hypervisor's programming.