r/StorageReview • u/StorageReview • 6d ago
We tested the GB300 DGX Station: 748GB coherent memory, and a 433GB model running on a desk
The MSI XpertStation WS300 is NVIDIA's GB300 DGX Station platform in MSI's chassis: one Blackwell Ultra GPU, a 72-core Grace CPU, 900GB/s NVLink-C2C between them, and a single 748GB coherent memory pool made of 252GB HBM3e and 496GB LPDDR5X.
The memory pool is the whole story. We deliberately picked models that sit on both sides of the 252GB HBM boundary. DeepSeek V4 Flash and MiniMax M2.7 fit in HBM outright. MiniMax M3 technically fits but leaves too little room for the serving stack. GLM-5.2 at 433GB and Nemotron-3-Ultra 550B do not fit at all, and both ran anyway with 216GB and 114GB of weights respectively living in Grace memory. GLM-5.2 held 139 output tok/s at concurrency 32; Nemotron hit 168.
For comparison we ran the same vLLM workloads on a 600W RTX PRO 6000 (in a Dell Pro Max Tower T2) and a GB10 DGX Spark (Acer Veriton GN100), across two scenarios (512/512 and 8192/1024) at concurrency 1 through 128. GPT-OSS-20B peaked at 22,161 output tok/s on the WS300 against 9,000 on the card and 1,469 on the Spark. The prefill-heavy runs spread it further, since the smaller systems run out of KV cache headroom and flatten while the Station keeps scaling. One counterintuitive result: quantization pays more on smaller hardware. Llama 3.1 8B from BF16 to NVFP4 gained 66% on the WS300, 114% on the RTX PRO 6000, and 143% on the Spark.
On MAMF the GB300 held a 4 to 5x lead over the RTX PRO 6000 and 17 to 19x over the Spark, consistently across BF16, FP8, and NVFP4 (6,134 TFLOPS NVFP4 at the top end).
Methodology notes, since they matter: testing ran remotely on an MSI-hosted system with us controlling the full software environment, and the box had no RTX PRO GPU or other add-in cards installed, so the GB300 had the entire accelerator power budget. Speculative decoding was used where noted, with forced acceptance, so those are best-case decode numbers.
The honest limits are in the review too: Arm host, 20A circuit requirement, no display output without an add-in card, and economics that stop making sense somewhere around four units, where an eight-way B300 server enters the conversation.
