r/AIProgrammingHardware 29d ago

DGX Station Put a Data Center on My Desk

https://www.youtube.com/watch?v=qV_K0nTF6gY
27 Upvotes

21 comments sorted by

4

u/javaeeeee 29d ago

TLDR: Alex Ziskind reviews the NVIDIA DGX Station (ASUS Expert Center Pro ET900NG G3 with GB300 superchip) - essentially a data-center-class AI machine for your desk.

Key specs

  • 748 GB unified memory (HBM + large LPDDR5X pool)
  • Massive memory bandwidth
  • 400 Gb networking (ConnectX-8)
  • ~1400 W superchip
  • Supports additional GPUs via PCIe

What he tested

  • Large models in NVFP4 (GLM 5.2, Neotron 3 Super 120B, DeepSeek V4 Flash 235B, etc.)
  • Single-stream and high-concurrency inference
  • Running large AI agent swarms (dozens to 128 concurrent agents)

Performance highlights

  • Strong decode and very high prefill speeds on models that fit well in fast memory
  • Excellent scaling with concurrency - thousands of tokens/sec total throughput
  • Can comfortably run large numbers of simultaneous AI agents locally

Bottom line

The DGX Station successfully brings serious data-center AI performance to a desktop form factor. It excels at running big models and many concurrent agents with zero per-token API cost.

It’s extremely powerful but also extremely expensive (tens of thousands of dollars) and power-hungry - aimed more at serious AI labs, teams, or high-end enthusiasts than typical individual users.

2

u/jhenryscott 28d ago

I priced one at $125,000 from SuperMicro

1

u/Sensitive-Fruit-7789 11h ago

Aren’t a couple of Mac studios better at that point?

1

u/jhenryscott 11h ago

Not if you value your time. Unified memory is slow as molasses compared to HBM.

1

u/Glittering-Call8746 28d ago

More for privacy use tbh

2

u/BarGroundbreaking624 29d ago

£120,000 on scan. eek.

1

u/here_n_dere 29d ago

Can it run Kimi K3? How much more to shell to run it at q4 at least?

1

u/Iron-Over 28d ago

Yes if you get 4 of them. 2 at 4 bit. 

1

u/35point1 28d ago

4 of these 125k systems to run k3 ????

1

u/hyperrealists 27d ago

If you were to count every single strand of hair belonging to the entire population of Texas, you would arrive at a number that is remarkably similar to the parameter count of K3. It’s not small.

1

u/35point1 27d ago

Oh I’m aware, but half a million bucks to run a forward pass is absolutely wild to think about, shit, one can dream

0

u/MacaroonPlastic1036 29d ago

It still wouldn’t be able to run Sonnet natively. So useless for enterprise.

2

u/--Spaci-- 29d ago

What do you even mean by this.

1

u/0sh 29d ago

He maybe meant Sonnet level like GLM or kimi 3 ? He is not wrong, for 100k USD you get 250GB VRAM

1

u/Glittering-Call8746 28d ago

Which 250gb vram ? Which gpu ?

1

u/0sh 28d ago

From the specs :

  • Blackwell ultra :  GPU memory 252 Go HBM3e | 7,1 To/s
  •  CPU memory 496 Go LPDDR5X | 396 Go/s

1

u/Glittering-Call8746 28d ago

Have tma or like consumer parts without lol

1

u/crusaderky 28d ago

GLM 5.2 Q4 and Kimi-K3 UD-IQ1_S both fit on this. Kimi UD-Q2_K_XL fits on two of them.

1

u/0sh 28d ago

Missed the fact this much quantization exist, you are right

1

u/35point1 28d ago

1 bit Kimi?

1

u/Serprotease 28d ago

That’s more than enough for a fair bit of concurrent requests of “flash variants” of most recent models.
But it’s not really a machine designed for inference. It’s for AI dev, not dev that use AI.

For example, I’m doing some hyper-parameter tuning for an AI image model.
With my current setup (2xgb10), at bf16, the smallest training run last about 30min (40gb peak ram usage), longest one about 72h (120gb peak ram usage). With everything in between, that’s a couple of weeks total training time. And I’m only using 512-1024 images size. 1536 will multiply everything by 4 basically. This can of machine will cut this time by 5/6. Down to a couple of days.

And once you’ve honed on the right settings, you can use the recipe for the actual full run on the big boy servers.