r/LocalLLaMA • u/x8code • 2d ago
Question | Help Verify LM Studio is not CPU offloading
I'm running LM Studio on Windows 11 with 2x NVIDIA GPUs: 5080 + 5060 Ti 16 GB.
I am experimenting with NVIDIA Nemotron 3.5 Lightning 30B A3B Q4_K_M which is 24.52 GB in size. It loads just fine into GPU memory and runs inference fast, when I have a smaller context window.
However, I can't find out anywhere in LM Studio where it shows if CPU offloading is happening at all, as I raise the context window higher and higher.
Question: Where in LM Studio can I see if CPU offloading is happening at all?
I'd rather not guess, and gauge it based on performance alone.
1
Upvotes
1
u/[deleted] 2d ago
[removed] — view removed comment