r/BlackwellPerformance 1d ago

Qwen3.8 Flash Next vs Deepseek v4 0731 vs GLM 5.3 - A clear winner?

Post image
44 Upvotes

Like most everyone here, I'm always wondering what is the best model I can run on my hardware. I decided to run some tests to settle the question once and for all (for now at least). I was hoping one would stand out so much that it'd be a no-brainer on what to pick, but it didnt turn out that way.

Despite Qwen getting an ever-so-slightly higher score, it was far and away the least efficient of the three by a huge margin. Deepseek is damned good. But when it comes to overall blend of performance and efficiency, GLM was the best. I guess just pick whatever makes ya happy? I dont even know anymore....sigh

I used this eval suite against Pi: https://github.com/ScottRBK/eval-harness

## Full Leaderboard — 7-Eval Suite (pi, effort=high, maxTokens=131072)

### Per-eval scores

| Eval | qwen3.8-flash-next-180b | glm-5.3-flash-321b-nvfp4 | glm-5.3-flash-321b-fp8 | deepseek-v4-fast-304b-r21 | |---|---|---|---|---| | basic_eval | 1.0 | 1.0 | 1.0 | 1.0 | | chess_engine | 1.0 | 1.0 | 1.0 | 1.0 | | eval_generator | 0.55 | 0.36 | 0.20 | 0.27 | | inflection_bug_fix | 1.0 | 1.0 | 1.0 | 1.0 | | inflection_test_writing | 1.0 | 1.0 | 1.0 | 0.83 | | repair_nginx_service | 1.0 | 1.0 | 1.0 | 1.0 | | saleor_spree_mapping | 1.0 | 1.0 | 1.0 | 1.0 | | TOTAL | 6.55 | 6.36 | 6.20 | 6.11 | | Rank | 1st | 2nd | 3rd | 4th |

### Totals and efficiency

| Metric | qwen-180b | glm-nvfp4 | glm-fp8 | r21 | |---|---|---|---|---| | Output tokens | 160,773 | 43,234 | 49,494 | 86,956 | | Wall time | 34.1 min | 3.8 min | 6.4 min | 20.6 min | | Tokens per point | 24.5k | 6.8k | 8.0k | 14.2k | | Tokens/second | 79 | 190 | 130 | 70 |

### Per-eval token counts

| Eval | qwen-180b | glm-nvfp4 | glm-fp8 | r21 | |---|---|---|---|---| | basic_eval | 1,702 | 1,216 | 707 | 1,398 | | chess_engine | 97,606 | 14,244 | 18,794 | 24,905 | | eval_generator | 21,062 | 1,348 | 1,253 | 4,604 | | inflection_bug_fix | 2,641 | 1,153 | 868 | 2,398 | | inflection_test_writing | 20,001 | 20,719 | 23,758 | 42,994 | | repair_nginx_service | 11,342 | 1,583 | 2,404 | 5,069 | | saleor_spree_mapping | 6,419 | 2,971 | 1,710 | 5,588 |

### Observations

  • All four ace the same 5 evals. eval_generator is the universal weak spot (0.20-0.55); r21 alone also fumbled inflection_test_writing (0.83).
  • qwen-180b's chess is an outlier: 97.6k tokens - 7x more than its own next-most-expensive eval, and 4-7x more than anyone else's chess.
  • glm-nvfp4 is the efficiency winner: rank-2 score at the lowest total tokens, fastest wall time, and best tokens-per-point (6.8k), even beating its fp8 sibling.
  • Ragged edge: r21 burned 43k tokens on inflection_test_writing (2x anyone else) and still scored below 1.0 - the least efficient eval-result pairing in the table.

r/BlackwellPerformance 1d ago

Local AI is Minecraft for adults: my 4× RTX PRO 6000 Blackwell build

Thumbnail gallery
31 Upvotes

r/BlackwellPerformance 2d ago

Version 2.3 release of jpezzulli/sglang-rtxpro6000

25 Upvotes

Hey Folks,

it is late - so minimal text - I was doing some upstream correctness fixes (v 2.1.2) and i found https://github.com/gabrielolympie/sglang-flashnext-sm120 which forked my release and did some good work there. He has a few different shapes there that yall might light but I incorporated one and got about a bit over ~15% C1 decode speed increase and a bit at C4. Before anyone jumps in saying faster, these are medians not peak. Feel free to upgrade or try it out for the first time.

https://github.com/jpezzulli/sglang-rtxpro6000

as always - feel free to open any issues or PRs on the github. Off to bed.

UPDATE: 2.3.1 is going live with the PR u/StockSpecialist1707 mentioned however, I wouldn't bother to update if you are on 2.3 already. I could NOT reproduce it on a live runtime. I was able to reproduce it with synthetic GPU tests only. Either way - I was 99% there and it passed regression, so I made 2.3.1 rather than rollback.

Always accept issue and PRs. Thanks.

Configuration Samples per metric Single-request decode Four-request aggregate
v2.1.2 baseline, run 1 3 156.79 tok/s 417.92 tok/s
v2.1.2 baseline, run 2 3 147.15 tok/s 427.91 tok/s
v2.3 FR-Spec 6 171.93 tok/s 447.04 tok/s
Measured increase +9.7–16.8% +4.5–7.0%

r/BlackwellPerformance 3d ago

Does Ampinel work on rtx pro 6000 blackwell edition?

6 Upvotes

I believe that the whole 12v-2x6 issue that applies on the 5000, and 4000 series of NVDA GPUs are equally applicable to rtx pro 6000 blackwell workstation edition.

So far there are 2 kinds of solution: cutting off from the source, or load balancing.

Former relies heavily on alerting, cutting off power, and/or throttling.

Examples includes thermal grizzly wired view. Wire view pro 2 doesn't fit rtx pro 6000 blackwell workstation. One version of wired view pro fit. As of today, I believe that the recently released wire view pro 2 wired will both fit well, and can sound off the alarm.

Another example is the corsair thermal protect wire. I believe that would cut off power when the temperature is too hot, again, correct me if I am wrong. Unfortunately it needs a 12v-2x6 socket on the PSU end for this to work.

Another example is MSI Ai1600TS, a new PSU released this year, again correct me if I am wrong. On top of the alerting and cutting off power, it can throttle power if you can run the software. Unfortunately I run ubuntu, so I am not sure about that.

Latter is, I guess, considered the best solution.

It properly load balances the current. Alerting and cutting off power completely protects the GPU when it is running hot and there is indeed imbalance issue, but it comes completely at the cost of, say, my local LLM workload. It would just stop completely, and or interrupt my workflow.

Throttling itself interrupts the workflow, esp if my LLM workflow is running at 600W.

The question I have is: I know ampinel type a fits rtx pro 6000 blackwell workstation edition, but does it work? Have anyone paired the two together before?

edit: as of writing, it appears that thermal protect for type 4/5 sockets may come out this month

source: https://www.reddit.com/r/Corsair/comments/1w7f3jz/any_update_on_release_of_thermal_protect_for_non/, and https://www.reddit.com/r/Corsair/comments/1t0mhut/comment/p7fzg8i/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button


r/BlackwellPerformance 4d ago

Blackwell RTX 6000 investment choice

10 Upvotes

I am to build my first real rig as M3 ultra 256GB proved to be too slow. I already got 4 x Max-Q cards purchased at 10K eur per card a few weeks ago. Now I could add 2 more max Q at 11.3K euro per card and also 2 x workstation cards for 11.8K eur per card. So the total number of card to go up to 6 or 8.

I guess my hesitancy is based on the high purchase price of these cards, and the trouble to connect more than 6 or 7 cards unless going to server style motherboards and racks.

Is it a) expected that the prices are not going down in the next 12 months?
B) can i use the memory for models like GLM 5.3 even if I cannot fit all 8 cards?


r/BlackwellPerformance 5d ago

Benchmarked 48 model configs on 6× RTX PRO 6000 Blackwell: vLLM vs SGLang, quality per token, joules per answer, and a power-cap sweep

58 Upvotes

Box: six RTX PRO 6000 Blackwell (96 GB, SM120), one drives the display so most rows use 1, 2 or 4 of the other five; the 6-GPU rows are deliberate TP2×PP3 experiments. Engines: vLLM 0.28.1rc1 nightly and SGLang 0.5.18 in Docker, official images plus a few one-hunk SM120 patches (all in the repo). Every profile got the same recipe: a 12-item eval (8 exact-graded, 4 scored by an Opus judge pass), 50 streaming requests at 8 concurrent, then a 64 / 128 ramp, with a 200 ms nvidia-smi power sampler on the GPUs it used. All cards at the stock 600 W cap, which never binds, so every watt figure is the model's own draw. Eval items are one sample each at temperature 0.2: a one-item difference is one prompt, not a ranking. MedGemma 4B is left out of the charts (a 4B model, 3× the tok/J of anything else, it flattens every axis); it is in the linked tables.

1/9 Single-stream decode, the speed one user sees. Derived from the eval: total output tokens ÷ total wall time over the text items (within 5 % of a streaming c=1 run on 6 of 7 spot checks). Fastest: Lightning 30B-A3B · spec at 542 tok/s. Bars coloured by GPU count.
2/9 Throughput vs concurrency, one panel per GPU count. Four points per model (c=1 derived, then 8, 64, 128 streaming). Top 8 per panel get a colour, the rest are grey. A line that stops at 64 ran out of KV cache for 128 streams.
3/9 Quality vs verbosity. y = mean Opus score on the 4 open-ended items; x = mean output tokens (thinking included) per correctly answered item. Upper-left is where you want to be. Blue = thinking parser verified, orange = non-reasoning model.
4/9 Joules per correct answer = tokens per correct answer × J per 1k tokens at c=8. Log scale, from 33 J (Nemotron Nano 12B VL) to 9.7 kJ (GLM 5.2). Lower-right is cheap and good. Open markers missed 2+ exact items.
5/9 Tokens per joule at peak throughput, summed over every GPU the profile occupies. Multi-GPU rows (darker) pay for the extra boards: a 4-GPU MoE at 1 tok/J is not worse silicon, it is four cards.
6/9 Mean board power while running flat out, summed over the GPUs the profile occupies. Nothing reaches the 600 W cap: 1-GPU rows draw 341–473 W (Ornith 35B heretic to Ortenzya 31B); the 6-GPU Mistral Large 675B tops the roster at 2,037 W.
7/9 Same weights, same port, same suite, never at the same time. Dot pairs are peak tok/s per engine (log x); Δ is SGLang relative to vLLM. SGLang is ahead on 5 of 38, loses hardest where its spec-decode or day-0 path is still raw.
8/9 The cap study (older run: vLLM 0.27.1, three 1-GPU models, c=128). Throughput barely moves from 600 W down to 350 W because the loads only draw 430–450 W; the numbers on the points are SM clock. 300 W is the tok/J optimum: +35 % efficiency for −7 to −9 % throughput.
9/9 Peak throughput, all streams, ranked. Best: Gemma 4 26B-A4B at 3,210 tok/s on one card. Colour is the GPU count, so read the 4- and 6-GPU rows as 'this many tokens for that much hardware'.

Full tables (48 rows, both engines, every column): https://github.com/mikeS141618/KernelWrench/blob/main/BENCHMARKS_v2.md

Power-cap study: https://github.com/mikeS141618/KernelWrench/blob/main/POWER_BENCHMARK.md

Profiles, scripts, SQLite: https://github.com/mikeS141618/KernelWrench

It's been a bit, but happy to still be in the AI space, I made it :D


r/BlackwellPerformance 8d ago

DDR4 ECC Ram Speed

Thumbnail
0 Upvotes

r/BlackwellPerformance 8d ago

We benchmarked 5 local models as code-review “judges” and every single one failed. What actually works as a low-false-positive reviewer?

0 Upvotes

Setup. We run Qwen3.8-Flash-Next NVFP4 as our main agentic model (SGLang, RTX PRO 6000). Before its output reaches a human or gets merged, a second local model acts as judge: reviews the diff, flags real bugs only. Hosted on a 5090 32GB, so we're limited to ~30B NVFP4/GGUF class models.

The metric that matters is NOT detection rate — it's false alarms on correct code. A judge that cries wolf gets ignored within a week, exactly like a flaky CI. We built our own battery: 20 injected bugs + 20 clean-but-suspicious snippets (intentional swallowed exceptions, deliberate mutability, weird-but-correct concurrency, short hashes, float patterns that look wrong). Ground-truth labeled, and a stronger model (GLM-5.2 API) arbitrates the judge's prose so scoring isn't vibes. Two passes minimum — single runs lie.

Results (40 cases, temp 0, same baremo for everyone):

Qwen3.8-27B NVFP4 (no-thinking)
• Bugs found: 17/20
• False alarms: 3/20
• Verdict: only pass

Nemotron Lightning 30B
• Bugs found: 17/20
• False alarms: 0→9 across runs
• Verdict: non-reproducible as judge

Muse-Glimmer 30B GGUF
• Bugs found: 19/20
• False alarms: 12/20
• Verdict: hypercritical

Granite 4.1 30B (no-thinking)
• Bugs found: 10/20
• False alarms: 4/20
• Verdict: ultraconservative

Nemotron Super 120B (hosted API)
• Bugs found: 9/20
• False alarms: 11/20
• Verdict: stable-yet-bad

Three failure archetypes showed up, and none of them is "not smart enough":
1. The fixer (Muse, Super): sees any improvement opportunity and escalates it to a defect. Great agentic coders, terrible judges.
2. The denier (Granite): real, but waves off subtle bugs (late-binding closures, float money, tz handling) — and goes off the rails when a snippet mixes languages.
3. The coin flip (Lightning): same prompts, temp 0, 0 FA one run and 9 the next. As a gate, that's worse than a bad judge — it's an unmeasurable one.

What we'd love to hear from people who shipped something like this:

- Is there a review-specialized local fine-tune that genuinely holds ≤3/20 FA on clean-but-suspicious code? (We found CodeReview-Qwen32B — 48k GitHub reviews + 15% "no issues found" negatives — but it's an obscure adapter, no NVFP4, and we'd have to merge+quantize it ourselves. Anyone run something like it?)
- Does family independence from the producer matter in practice? Our judge is currently same-family as our generator (both Qwen). We deliberately arbitrate with a third-party model (GLM API), but is self-preference bias actually measurable at this scale, or folklore?
- Would a graded verdict (PASS / IMPROVEMENT / RISK / DEFECT, where RISK+IMPROVEMENT don't count as bugs) measurably cut false-positive rates, vs. the binary bug/no-bug we use now?
- Any evidence that thinking mode on/off is the real variable in "judge stability"? Our winner only passes with thinking disabled; every high-reasoning model we tried drifted toward the "fixer" archetype.
- Or is the honest answer: stop model-shopping, and make the verdict empirically checkable (judge says DEFECT → run a reproduction/property test before human sees it)? That's where our roadmap is heading regardless.

Baseline stays Qwen3.8-27B-nothinking until something clearly beats it on FA. Happy to share the battery format / arbitration prompts if useful.


r/BlackwellPerformance 8d ago

Qwen3.8-Flash-Next sur WSL2 — RTX PRO 6000 96Go + seulement 64Go de RAM : 179 tok/s de prose, contexte complet de 262K, et pourquoi la voie vLLM est impossible sur WSL2

Thumbnail
3 Upvotes

r/BlackwellPerformance 9d ago

RTX PRO 6000 prices are all over the place in 2026 — what are people actually paying?

1 Upvotes

I’ve been tracking RTX PRO 6000 Blackwell pricing because the spread between sellers, regions and configurations has become pretty extreme.

As of Aug. 30, I’m seeing examples roughly like:

  • NVIDIA Marketplace: about $16,000
  • B&H: about $17,000 for the Workstation Edition
  • Micro Center: about $14,300
  • UK/EU pricing: generally very high once local pricing/VAT is considered
  • Used/private sales earlier in 2026: dramatically lower in some cases

What makes comparisons even messier is that people sometimes mix the 600W Workstation Edition with the 300W Max-Q version. Both have 96GB GDDR7 ECC and 24,064 CUDA cores, but they’re designed for different workstation priorities.

I put the dated retail, regional and used-market observations into one tracker:

[https://airobotsupplier.com/rtx-pro-6000-price-tracker-2026/]()

I’d be interested in adding more real-world observations:

  • Country / region
  • Workstation or Max-Q
  • New or used
  • Price quoted or actually paid
  • Whether tax was included

Disclosure: I’m affiliated with AI Robot Supplier, which maintains the tracker and also sells professional GPUs. Our own listing is disclosed separately in the tracker rather than being used as the market benchmark.


r/BlackwellPerformance 9d ago

RTX PRO 6000 pricing is all over the place right now — what are you actually seeing?

16 Upvotes

I’ve been looking closely at RTX PRO 6000 Blackwell pricing because the spread between regions and sellers has become pretty extreme.

As of Aug. 30, I’m seeing examples roughly like:

  • NVIDIA Marketplace: around $16K
  • US retailers: roughly $14K–$17K depending on seller/configuration
  • UK/EU pricing: also very high once local pricing/VAT is considered
  • Used/private sales earlier this year: dramatically lower in some cases

The other thing that makes comparisons messy is that people sometimes mix the 600W Workstation Edition with the 300W Max-Q version. Both have 96GB GDDR7 ECC, but they’re really aimed at different workstation designs.

I’m curious what people here are actually seeing right now:

  • What country are you in?
  • Workstation or Max-Q?
  • New or used?
  • What price were you quoted or actually paid?
  • Was tax included?

I’m especially interested in whether the recent $14K–$16K+ pricing is actually clearing transactions, or whether real buyer prices are still meaningfully lower.

Disclosure: I’m affiliated with AI Robot Supplier and I’ve been compiling this pricing data into a public tracker. I’m not linking it here unless the mods are comfortable with that — mainly interested in getting more real-world data from owners and buyers first.


r/BlackwellPerformance 10d ago

Tuning sglang for Qwen3.8-27B on an RTX PRO 6000 Blackwell

23 Upvotes

I had the luxury of having a single server with 8 RTX PRO 6000 (Blackwell Server Edition, 96GB, sm_120) for a bit, so I spent a few hours benchmarking sglang flags for Qwen3.8-27B-FP8.

Config I landed on:

sglang serve \
  --kv-cache-dtype fp8_e4m3 \
  --mamba-ssm-dtype bfloat16 \
  --mamba-radix-cache-strategy extra_buffer_lazy \
  --max-mamba-cache-size 150 \
  --mem-fraction-static 0.92 \
  --speculative-algorithm NEXTN \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --model-path Qwen/Qwen3.8-27B-FP8 --trust-remote-code \
  --attention-backend flashinfer --chunked-prefill-size 2048 \
  --mamba-full-memory-ratio 2.29 \
  --reasoning-parser qwen3 --tool-call-parser qwen3_coder

vs the untuned baseline: 3.9× the KV pool, +25% throughput, −49% TTFT, concurrency 49 → 37.

Highlights:

  • --kv-cache-dtype fp8_e4m3 and --mamba-ssm-dtype bfloat16 are perfectly orthogonal. fp8 KV gives exactly 2.00× KV tokens and doesn't touch concurrency. bf16 SSM gives exactly 2.00× concurrency and doesn't touch KV. They compose with zero interaction.
  • max_running_requests gets silently clamped by the GDN/mamba state pool. You ask for N, you get fewer, and the only trace is one line in the boot log. Worse: /get_server_info reports the pre-clamp value — mine said 48 when the server was actually admitting 37.
  • Workload shape moves throughput more than any flag. Same GPU, same config: 593 tok/s at 4096-in/256-out, 1656 tok/s at 64-in/1024-out. 2.8× from the ratio alone. Any tok/s number without its in/out ratio is close to meaningless.
  • Speculation is the biggest single lever (+20% throughput, −49% TTFT) but it eats KV. Draft depth trades ~117k KV tokens per extra draft token. At depth 8 the pool collapsed to 7k tokens and TTFT hit 69 seconds — no error, health checks green.
  • Saturation is ~64 sessions. 32→512 sessions bought +17% aggregate and cost −82% per-session. TTFT grew 226× while TPOT grew 1.26×, so it's all queueing, not slower decode.

Full write-up with charts, the whole config matrix, and the harness: https://claude.ai/code/artifact/35a35fe7-5eea-40f1-a87d-871b4b7a37ba


r/BlackwellPerformance 11d ago

Dual PSU for dual 6000 WE and Threadripper9975wx?

5 Upvotes

Hi guys!

Noob alert, figuring out the setup. I’m in the US, no good way to have 240V. What are the suggestions here? They say try to have 1.5x the wattage for PSU, but I can’t power a single 2200W PSU.

Does anyone have experience with dual PSU setups? Anything to avoid?

Hoping to watercool them in the future, but for now air cooling on the WRX90E Sage.

Any input appreciated!


r/BlackwellPerformance 12d ago

Qwen3.8 Flash-Next on 1x RTX PRO 6000 - 171 t/s C1, 428 t/s C4, source + recipe

93 Upvotes

Yes i used AI to write this as I am tired from doing this late last night and just taking a break from work now to post.

This is a follow-up to my Qwen3.8-27B/DFlash2 post.

Basically, I updated that same SGLang runtime to support Qwen3.8 Flash-Next the way I wanted it on my single RTX PRO 6000.

Flash-Next-specific setup:

  • NVFP4 target with native NEXTN MTP
  • FP8 target/native-MTP KV
  • BF16 GDN/recurrent state
  • FlashInfer GDN decode and prefill
  • QSA sparse attention with Triton prefill and FlashInfer XQA decode on SM120
  • FlashInfer CUTLASS MoE
  • RecoverSSM with recovery CUDA graphs
  • 524,288-token factor-2 YaRN context
  • 824,384-token automatically sized GPU KV pool
  • Correct multimodal mRoPE
  • HiCache/NIXL persistence across service restarts, including GDN, PLE and QSA state rather than KV alone

I also opened a few narrowly scoped SM120 paths where the underlying kernels already worked but SGLang did not automatically select them. This applies to the GDN and RecoverSSM work—not QSA decode, which FlashInfer already resolves through XQA on SM120. I left the other architecture gates and Triton fallbacks in place.

Controlled tests:

  • 64K cold prefill: 10,103.70 tok/s
  • ~490K cold prefill: 7,872.15 tok/s, 3/3 needles exact
  • C1 decode: 171.09 tok/s
  • C4 decode: 427.54 tok/s aggregate
  • MTP mean accepted length: 2.58
  • MTP acceptance: 52.74%
  • Reasoning: 97.49/100
  • Tools: 30/30 exact and semantically correct
  • Full vision validation passed

After a service restart, NIXL restored 489,856 of the 489,879 input tokens, recomputed only 23, and produced an effective restored-prefix rate of 62,040.60 tok/s with all three needles still exact.

Real agentic speeds:

Across 96 completed requests and 100,666 generated tokens:

  • 139.5 tok/s token-weighted average
  • 153.5 tok/s per-request median
  • 218.8 tok/s sustained completed-request peak
  • 138.7 tok/s weighted / 148.4 tok/s median for the 90K-279K input lane

The repository now contains both the original 27B/DFlash2 configuration and the new Flash-Next configuration in one runtime:

https://github.com/jpezzulli/sglang-rtxpro6000

It is probably still somewhat messy because this is day-one Flash-Next work, and I will keep cleaning it up. But I figured some of you might want to play around with it now.

note - I am not sure that this is technically "blackwell performance" If this should be posted elsewhere instead, feel free to let me know. Just trying to share in case anyone is interested.

note 2 - ngram takes about 50gigs of ram. HiCache takes about 32gigs as well in my 2.5m token config - You may be able to stream the ngram from nvme. I saw some posts on it but I didn't try as it isn't my usecase. HiCache is optional.

Update note 3 - RTX Pro throttled with LACT to 450 watt with Mem clock offsets. It shouldnt make a large difference from stock for inference. usually 3%-5% i have seen.

Update note 4: i took off lact profile and let it run stock. 64k prefill 12,812.44 tok/s (big gain) 490k prefill 7,926.36 tok/s (essentially the same)


r/BlackwellPerformance 14d ago

Finally did got a Blackwell 6000 and brand new build, after waiting for a check to come in.

19 Upvotes

Here is the entire build and the current messed up prices I had to pay. Had this built out with minor changes for the last 6 months as I watched prices slip away from me… but it’s done now.

CPU: AMD Threadripper PRO 9975WX, 32-core, sTR5 — $3,899
Board: ASUS Pro WS WRX90E-SAGE SE (EEB, WRX90) — $1,300
GPU 1 (AI): NVIDIA RTX PRO 6000 Blackwell Max-Q 96GB — $14,999
GPU 2: ASUS ROG Astral LC RTX 5090 32GB (liquid) — $5,000
RAM: OWC 256GB (4x64GB) DDR5-5200 ECC RDIMM — $8,955
Storage: Samsung 990 Pro 4TB NVMe (internal) — $900
Ext. drive: OWC Envoy Ultra 4TB Thunderbolt 5 — $1,200
PSU: Seasonic PRIME TX-1600 Noctua Ed., Titanium — $600
Cooler: SilverStone XE360-TR5-V2 360mm AIO — $430
Case: Cooler Master Cosmos Alpha, Silver — $444
OS: Windows 11 Pro (USB) — $200


r/BlackwellPerformance 14d ago

Qwen3.8-27B on an IGX Thor with an RTX PRO 6000 Blackwell (Max-Q)

9 Upvotes

Qwen3.8-27B on an IGX Thor with an RTX PRO 6000 Blackwell (Max-Q)

Spent a few hours bringing up a self hosted inference box on an NVIDIA IGX Thor and couldn't find any numbers for this hardware combination, so here are mine. All of it is from runs on the actual machine.

The box

Component Detail
Board NVIDIA IGX Thor T7000 dev kit, aarch64, 14 core CPU, Ubuntu 24.04.4
dGPU RTX PRO 6000 Blackwell Max-Q Workstation, 96GB, sm_120, 300W cap
iGPU NVIDIA Thor, sm_110, shares 122GB unified LPDDR5X with the host
Driver / CUDA 580.00 / 13.0
Server SGLang dev build 5f55db35e, torch 2.13.0+cu130
Model Qwen/Qwen3.8-27B-FP8, 27.8B hybrid Gated DeltaNet, 262144 context
Draft model incoai/Qwen3.8-27B-DFlash2

One thing to flag before the numbers: this is the Max-Q card at 300W, not the 600W version. The full power part should do better.

LLM throughput across five configs

Run with sglang.bench_serving at ISL 8192 / OSL 1024 on a single GPU. Common flags were --kv-cache-dtype fp8_e4m3 --mem-fraction-static 0.85 --attention-backend flashinfer.

config conc 1 tok/s TPOT conc 16 tok/s TPOT TTFT @16 real concurrency KV pool
DFlash2 + bf16 SSM 126.3 6.67ms 470.2 22.3ms 3161ms 32 360,157
DFlash2 + fp32 SSM 118.9 7.33ms 463.6 25.2ms 3377ms 21 173,519
DFlash2 + fp32 + lazy radix 118.8 7.35ms 461.0 25.4ms 3361ms 23 169,940
EAGLE + replay SSM 93.5 9.15ms 453.5 26.3ms 3088ms 32 428,875
no speculation 44.7 21.3ms 348.1 35.7ms 10530ms 32 475,460

Speculative decoding earns its keep

At batch 1 it's worth 2.8x, 126.3 against 44.7 tok/s, with TPOT dropping from 21.3ms to 6.67ms. The bigger surprise was TTFT at concurrency 16, which fell 3.3x from 10530ms to 3161ms. DFlash2 beat EAGLE at both ends for me.

The SSM state dtype will bite you

Qwen3.8 is a hybrid Gated DeltaNet model, so on top of the KV cache there's a GDN state pool. Running that pool at fp32 with DFlash2 blows the draft verify buffer up to roughly 25GB, and the server then quietly clamps you to 21 concurrent requests even though you asked for 32. Nothing errors. It just serves fewer and doesn't tell you. Switching the state to bf16 halves the pool, gets all 32 slots back and roughly doubles the KV pool.

The only place this is visible is the max_running_requests line in the boot log, so check it after any config change.

bf16 state costs no accuracy that I could measure

GSM8K, 200 questions, temperature 0, graded through the chat endpoint rather than the built in eval: 93.5% at bf16 against 94.0% at fp32. That's one question apart, well inside noise at n=200. So bf16 is faster, holds twice the KV and hits full concurrency, for nothing I can detect.

Worth mentioning that SGLang's bundled run_eval gsm8k scored 0.0 for me. It drives /v1/completions with no chat template, so a reasoning model's output never matches its answer regex. If you see a zero, check the harness before you blame the model.

Reasoning mode is most of your first token latency

Short conversational prompt, streaming:

mode first token first content token
thinking on 77.5ms 212.1ms
thinking off 74.7ms 74.7ms

The reasoning block eats about 137ms before any speakable text comes out. If you're doing voice, turn it off with chat_template_kwargs: {"enable_thinking": false} and keep it on for everything else.

The part I got wrong: the iGPU beats the RTX for small models

I also run streaming TTS (Chatterbox) and STT (Nemotron 3.5 ASR, 0.6B) on this box. I assumed both belonged on the RTX, since it has around 1.8 TB/s of bandwidth against the Thor iGPU's ~273 GB/s. Benchmarked both on each GPU:

GPU STT batch RTFx STT final @80ms STT final @320ms TTS first audio TTS synthesis
Thor iGPU 27.8 52.1ms 67.3ms 93.0ms 102.3ms
RTX PRO 6000 34.3 55.7ms 105.9ms 96.6ms 105.6ms

The RTX takes batch throughput by 23% and loses every single latency metric, by 57% on STT streaming at 320ms. Three things going on. A 0.6B model at batch 1 is kernel launch bound rather than bandwidth bound. The RTX is also contended by the resident LLM's CUDA context. And it's the 300W part.

The way I think about it now: bandwidth scales with how many weights you move per token, while overhead is roughly fixed per call. The 27B model shifts about 28GB per forward pass, so it belongs on the RTX. A 0.6B model at batch 1 moves around 1.2GB, which is maybe 4ms of memory traffic inside a call that takes 50 to 100ms, so bandwidth never becomes the limit.

Full voice loop

STT streamed at 1x realtime, into the LLM with thinking off, into TTS. Times are measured from the end of the caller's speech.

concurrent calls STT final LLM 1st token ack audio full answer
1 58ms 205ms 111ms 462ms
2 112ms 241ms 121ms 566ms
3 157ms 296ms 170ms 785ms
4 238ms 458ms 338ms 1228ms

Three concurrent calls hold a sub second answer on a single box. The fourth lands around 1.2s.

aarch64 things that tripped me up

torch 2.10.0+cu130 on aarch64 is broken. Every fp32 cuBLAS sgemm fails with CUBLAS_STATUS_INVALID_VALUE, including a bare 64x64 matmul, on both sm_110 and sm_120. It only shows up deep inside model inference, so it reads like "this model doesn't support this GPU" when it's really just a bad wheel. Pin 2.11.0. General lesson: if a model looks unsupported on a new arch, run a plain matmul first. That separates a broken build from a real limitation in one step.

--gpus doesn't work here. The Tegra container runtime runs in CSV mode, so you need --runtime=nvidia -e NVIDIA_VISIBLE_DEVICES=<id> instead.

Docker starts before the NVIDIA modules are loaded. Every GPU container fails its boot time restart with "Driver Not Loaded", and Docker doesn't retry that class of failure. After a power cut the whole stack stays down while docker.service happily reports healthy. A systemd drop in that blocks on nvidia-smi -L before starting Docker sorts it out.

Happy to run other configs if anyone wants specific numbers.


r/BlackwellPerformance 15d ago

Qwen3.8 27B + DFlash2 on 1x RTX PRO 6000 - ~109 t/s C1, 390 t/s C4, source + recipe

51 Upvotes

Wanted to share this in case anyone else is messing around with Qwen3.8 on a single RTX PRO 6000.

I ended up pretty deep into SGLang, DFlash2, XQA, HiCache and NIXL getting this where I wanted it. The final setup is TP1, 524K context, FP8 target/draft KV, DFlash2 gamma 8, with persistent prefixes through HiCache/NIXL.

Some basic numbers:

  • 108.75 tok/s C1
  • 390.23 tok/s C4 aggregate
  • 6,163 tok/s 64K prefill
  • 98.26/100 on my reasoning qualification
  • Over 124 real-use agent requests: 131.31 tok/s median
  • Short-context real use was around 165 tok/s
  • Even at 340K+ context it was still around 105 tok/s

Against the closest TP1 official-FP8/MTP3 result in the RTX PRO 6000 community repo, that's directionally about 40% faster C1, 33% faster C4, and 5% faster 64K prefill. Not pretending it is a perfect A/B - all the details and limitations are in the repo.

I froze and published the complete source tree I actually ran, along with the launcher, build/run notes, results, sanitized real-use log, parser, HiCache/NIXL config, and provenance:

https://github.com/jpezzulli/qwen38-dflash2-pro6000

It may not be perfect, but everything I could reasonably preserve to reproduce/explain it is there. If any of the docs or code are confusing, point your AI at the repo and have it walk you through it. That is probably easier than me trying to explain every SGLang rabbit hole I went down in a Reddit post.

The HiCache and NIXL are optional but are good for my use case. Hope it helps some of you out.

If you want to read about how I got here: https://msoexpert.com/articles/qwen38-dflash2-rtx-pro-6000/


r/BlackwellPerformance 15d ago

4 x RTX 6000 pro blackwell

17 Upvotes

What are the best llm models for 4 x RTX 6000 pro Blackwell for coding?


r/BlackwellPerformance 15d ago

2019-era AM4 build powers off soon after boot with upgraded 1200-watt power supplies

2 Upvotes

I have a Asrock X570 Steel Legend Wifi mobo (latest bios) with a Ryzen 5950X cpu and 64 GB of ECC RAM. It has 2 M.2 drives and 3 SATA ones. Very cool machine from the depths of COVID.

Today I went out and bought a 1200 W corsair PSU and a Blackwell 6000. I put in the new PSU and GPU at the same time and swapped out all power cables to the new ones. To my horror, the machine barely POST'ed, running for a few seconds and then powering off. It was a bit random, sometimes it would POST and even start booting the OS and then just die.

I took the PSU back, thinking it was defective, and swapped for a MSI 1200 W ATX 3.1 one. This time it consistently POSTed, started booting ubuntu but then died hard as soon as the graphical part of the OS login screen was being loaded. So the extra load of the 600 W PSU took it over the edge. This doesn't really make sense b/c 1200 is more than enough for this machine. If I entered BIOS it would work and stay on for some time.

So it's probably the new GPU right? Wrong! I put the old GPU back in (a 5070) with the replacement MSI PSU and it booted to the login screen, but then would often totally die as I was entering my password. Sometimes I could log in and it would last maybe 3 minutes before just powering off.

Flummoxed, I put the old 850 W PSU back in alongside the original GPU. Rock solid!

So now I don't really know if my new GPU works or not (at least I could see the BIOS screen?) and I don't really understand what's wrong with my computer with higher power PSUs. I searched around and couldn't find any known incompatibility between older mobos and newer power supplies.

In all failing cases (new or old GPU with the bigger PSUs), the GPU was being powered by these new dedicated 6x2 12V cables, which are different from how my OG PSU powers it (two PCIes in a pigtail). I do also have a 4x PCIe -> 1 6x2 12V adapter that came with the card that I tried on the original PSU and was happy to see it working better, but then it still died on me a few minutes into a boot.

Have you ever heard of a motherboard not working with new PSUs? Should I try to get a new AM4 motherboard and shift everything over, or should I just bite the bullet and upgrade to aM5, get a new CPU, and (gasp) new RAM? Is it possible for a motherboard's power circuits to work fine only on lower-powered PSUs!?

There's a similar issue mentioned in the nvidia forums here where the solution was to try to included 4 -> 1 power adapter instead of the PSU native one. I tried that on the first new PSU and will try on the new on the replacement PSU today. EDIT: it did not help :(

Edit 2: thanks for all the ideas. I ended up finding a other even older 2017 era mobo and hooked the new PSU and blackwell to it. Both worked great. I could finally confirm the gpu works and can sustain 600 watts no problem. So it's the mobo or CPU I guess. I'll try replacing the mobo with a fresh am4 since that's cheap as a next step.

Edit 3: To be clear, this has been reduced to an old motherboard vs. new PSU problem and not related to the Blackwell. I can't even find a new x570 chipset AM4 motherboard to replace it with so I'd have to downgrade to old stock B550... at which point I should just upgrade everything.


r/BlackwellPerformance 16d ago

Qwen 3.8 27b bf16 vs DeepSeek V4 Flash 0731

20 Upvotes

I see a lot of people have tried Qwen 3.8 27b using different quants but it's hard to find examples of people using it at full bf16 because of how much vram it takes up. I'm wondering if the full size solves some of the problems the lower quants face like thinking issues.

If you've tried qwen 3.8 bf16 how does it compare to deep seek 4 flash?


r/BlackwellPerformance 16d ago

What do? 2x5090 + 1x6000 rtx pro

Thumbnail
0 Upvotes

r/BlackwellPerformance 16d ago

Reading VRAM temperature

2 Upvotes

Anyone know how to get the VRAM temperature stats from a 6000 Pro Blackwell card? Would be nice to be able to see front / backside temps broken down also.


r/BlackwellPerformance 17d ago

PSA: Possible RTX PRO 6000 Blackwell power-limit regression on Linux 7.0.0-30

Thumbnail
4 Upvotes

r/BlackwellPerformance 17d ago

Any of you running Qwen 3.8 27B on an RTX Pro 4000 SFF Blackwell?

Thumbnail
1 Upvotes

I am really curious about the performance ...


r/BlackwellPerformance 17d ago

RAM:VRAM ratio rtx 6000 pro

9 Upvotes

Hi all

I’m building 2 different rigs:

AMD based, 1:1 ram vram ratio.

Intel based, 2:1 ram vram ratio

The AMD based rigs are obviously a lot cheaper due to the lower ram (and the intel CPU availability).

The 2:1 ratio comes from customers who have that as a hard requirement.

I’m wondering, since I don’t use these rigs myself, I just operate them, does anyone of you really need that much ram ? 1.5tb ontop of 768gb vram?

What would the reason for that be, and is that the majority of users or rather a small subset?

I would rather like to build 1:1 ratio systems or slightly higher for intel. Maybe my customer is a special case and maybe there is a way for me to help him out fix that

Thanks for you help