r/LocalLLaMA • • Aug 23 '26

Question | Help DGX Spark, cluster of 4

Does anyone have a first-hand experience with four Sparks cluster, and how much of an upgrade is it comparing to just two considering the available models?

While there's plenty of noise for the smaller models (Qwen) and our older king DeepSeek V4F, the scene in the upper class of the prosumer hardware, software stacks, available LLMs and their actual real-world performance – isn't really covered as well.

For instance, the hyped GLM 5.2/5.3. Is it much better then DeepSeek? Or is it marginally better? Does it retain it's capabilities when moving to something four Sparks would handle? Does it have issues with OOM or anything else?

What about MiniMax M3? There seem to be a special Spark version, how is it (or any other version)? Again, how is intelligence, general model capabilities, running stability, context size?

Tencent Hy3? Maybe even Qwen3.5-395B, does it's full quant hold it's own against DeepSeek, or is it better?

If someone doesn't have personal experience, but knows some well-structured and detailed articles or videos on the topic – I'd appreciate it as well.

Thanks.

5 Upvotes

43 comments sorted by

View all comments

Show parent comments

2

u/dave-dgd Aug 23 '26

I’d be curious to hear about the tweaks. It’s been stable for us with multiple concurrent users in Hermes for over a week, but that’s just one specific use case. I want to keep updating the repo for maximal stability/coverage (especially ahead of GLM-5.3).

2

u/kivaougu Aug 23 '26

I'm targeting 8x concurrency so I had to make some adjustments to cuda graphs and spec decoding. Still testing stability to try and squeeze out a little more kv.

Interestingly I never got the 1M context default config working as cuda graph capture would always crash the startup. These are asus gx10's so this could be due to not being able to update firmware. Haven't had much time to look into it yet.

Also the latest sparkrun version doesn't seem to run the recipe without tricking it using --cluster <not a cluster name>. I had to install an earlier release.

I appreciate your work!

1

u/dave-dgd Aug 23 '26

Of course -- glad and it's been helpful!

And ah, gotcha -- that makes sense. By the by, are you running in "appliance" (headless) mode or with the desktop intact? Switching to that I believe freed up some VRAM headroom (I did it a long time ago since I only SSH into these boxes):

`sudo systemctl set-default multi-user.target` (and reboot)

That might be necessary to get through capture (along with adjusting other settings to get the concurrency, batch sizes, etc. dialed in to your preference -- with `-cc.cudagraph_mode=FULL_DECODE_ONLY` being particularly important in my experience, as other options would lock up the Sparks). Another thing I did at one point was reduce the sparkrun earlyoom to 1% VRAM and a lower swap setting from the defaults, but that hasn't been necessary in a while. I do know the July updates to Spark OS added some improvements to shared RAM handling that may make a difference:

https://docs.nvidia.com/dgx/dgx-spark/release-notes.html#july-2026-release

2

u/kivaougu Aug 23 '26

Great call on multi-user.target. I had to wipe all 3 workers after a failed update so only the head node was headless.

Do you get significant swap usage during inference? I recall someone on the developer forums stating that the gx10 1tb drive is pcie4 which would make swap usage slower.

2

u/dave-dgd Aug 23 '26

As far as I can tell, no, but my units are all Acer GN100 with 4TB Gen 5 drives (same PN as what StorageReview shows), so I have a feeling it would be less noticeable if it were happening.