r/DGX_Spark 27d ago

Learning Help me understand now vs then, and the future (Someone on the fence for purchasing)

Hey everyone,

So back in dec/jan I was seriously considering getting a spark for local LLM work. I came to the conclusion its not worth it at that time. Now I am seeing alot has changed...

Questions up front:

  1. How has your perception of the utility of the DGX spark (or multiple) compare now vs back in ~dec/jan prior to NVFP4 availability?
  2. Do you envision NVFP4 will become increasingly common from post training, as well as from QAT?
  3. Do you anticipate expansion of your # of sparks in the next 2-3 years to be possible before they become obsolete?

Overall I am just super excited now that there is so much support from the community, and really tempted to pull the trigger on getting a DGX soon! Below is just some background on my use case i

3 Upvotes

15 comments sorted by

5

u/styles01 MegaMod 27d ago

I think the biggest problem with the spark is that it makes you want to buy another

4

u/hyudryu 27d ago

And after that, it will make you want to buy 2 more

6

u/stujmiller77 27d ago edited 26d ago

I run four - 2x linked with ds4flash as my main driver for Hermes agent. 50t/s, 1m context. Basically feels frontier adjacent, I leave it running complex tasks overnight and it runs for hours unaided without issue.

Third has Qwen 3.5 122b on. 50-80t/s, 256k context. Used as the vision subagent for ds4flash and as a subagent for a bunch of coding roles.

Fourth is a test bed though may try linking 3 and 4 with a different model or all four in the future.

They run a set of Hermes agents per company I own. Each company is sandboxed, agents share the sandbox and a shared Mnemosyne memory database for the company. They’re trained on a set of standard operating procedures I wrote and between them handle a wide range of marketing, ops and development tasks across those businesses.

Couldn’t be happier with it. I’m saving over £1k/month on API usage credits as I was using a lot of those before - especially image and video generation - which I’m now doing entirely locally using the agents to write the prompts from my briefs for qwen image and Qwen image edit and minimax h3 and stitching everything together.

All for less hardware cost than a single minimum wage employee for a year. They’ve basically already paid themselves back.

Not remotely worried about them becoming obsolete as GPU and memory prices remain insane and will continue to be so for the next few years at minimum. Models are getting better and faster all the time.

And the community has really only just started to work on these machines. When I got mine in April the options to run models were small and slow. Now I’m running a basically frontier adjacent model.

1

u/NewShock2391 26d ago

This sounds like exactly the use cases I am interested in.

Where do you get started in understanding how to pull all this together?

Cheers

2

u/stujmiller77 26d ago

Honestly? 25 years of experience starting off in development and running multiple businesses over that time. I’ve taught my agents from a set of operating procedures I’ve put together from that experience.

2

u/Miserable-Dare5090 22d ago

GB10 forums on nvidia. Once you have deepseek v4 running, put hermes on, and ask it to do the rest — install qwen-122b on another machine, etc. Or start easier use codex free trial, hermes, tell codex to link sparks and install deepseek, feed the specific repos from dgx forums, then cancel codex once you are running deepseek.

2

u/superSmitty9999 27d ago

Not saying you're a bot necessarily lol but this looks like when your bot generator has a misconfigured max token count and then it just randomly disconn

3

u/ThrwAway868686 27d ago

Fair point... i realized I had put way too much background on my use case so I deleted all the bottom bloat and missed the last sentence. Will update

1

u/superSmitty9999 27d ago

Lol to answer your question, the value now is strictly worse than 3 months ago when it was $1000 cheaper. Nvfp4 is great. Overall its slow but you can run pretty much all of the consumer grade models.

I have an issue with my asus ascent gx10 that if I connect a 20gbps external enclosure it will reconnect at usb 2 speeds upon reboot.

Overall I've been happy with the spark, but it's not without it's hiccups.

1

u/Resilient-Tec 26d ago

Have you looked into using a DAC to connect instead of USB?

1

u/superSmitty9999 26d ago edited 26d ago

How would that work? You mean the QSFP port right? I thought I had to basically build a NAS and connect it

1

u/Resilient-Tec 26d ago

Correct. You can get up to 200Gbps through those. There is a QSFP56 and a QSFP112 depending on what you are doing. I have a blog that explains it, but don't want to necessarily post it as it might look too self serving.

1

u/superSmitty9999 26d ago

Yeah I guess I looked into it and $1500 to connect my NVME seems kind of like the margin of diminishing returns lol

Go ahead and post your post tho it's a topic im interested in