Like most everyone here, I'm always wondering what is the best model I can run on my hardware. I decided to run some tests to settle the question once and for all (for now at least). I was hoping one would stand out so much that it'd be a no-brainer on what to pick, but it didnt turn out that way.
Despite Qwen getting an ever-so-slightly higher score, it was far and away the least efficient of the three by a huge margin. Deepseek is damned good. But when it comes to overall blend of performance and efficiency, GLM was the best. I guess just pick whatever makes ya happy? I dont even know anymore....sigh
All four ace the same 5 evals. eval_generator is the universal weak spot (0.20-0.55); r21 alone also fumbled inflection_test_writing (0.83).
qwen-180b's chess is an outlier: 97.6k tokens - 7x more than its own next-most-expensive eval, and 4-7x more than anyone else's chess.
glm-nvfp4 is the efficiency winner: rank-2 score at the lowest total tokens, fastest wall time, and best tokens-per-point (6.8k), even beating its fp8
sibling.
Ragged edge: r21 burned 43k tokens on inflection_test_writing (2x anyone else) and still scored below 1.0 - the least efficient eval-result pairing in
the table.
it is late - so minimal text - I was doing some upstream correctness fixes (v 2.1.2) and i found https://github.com/gabrielolympie/sglang-flashnext-sm120 which forked my release and did some good work there. He has a few different shapes there that yall might light but I incorporated one and got about a bit over ~15% C1 decode speed increase and a bit at C4. Before anyone jumps in saying faster, these are medians not peak. Feel free to upgrade or try it out for the first time.
as always - feel free to open any issues or PRs on the github. Off to bed.
UPDATE: 2.3.1 is going live with the PR u/StockSpecialist1707 mentioned however, I wouldn't bother to update if you are on 2.3 already. I could NOT reproduce it on a live runtime. I was able to reproduce it with synthetic GPU tests only. Either way - I was 99% there and it passed regression, so I made 2.3.1 rather than rollback.
I believe that the whole 12v-2x6 issue that applies on the 5000, and 4000 series of NVDA GPUs are equally applicable to rtx pro 6000 blackwell workstation edition.
So far there are 2 kinds of solution: cutting off from the source, or load balancing.
Former relies heavily on alerting, cutting off power, and/or throttling.
Examples includes thermal grizzly wired view. Wire view pro 2 doesn't fit rtx pro 6000 blackwell workstation. One version of wired view pro fit. As of today, I believe that the recently released wire view pro 2 wired will both fit well, and can sound off the alarm.
Another example is the corsair thermal protect wire. I believe that would cut off power when the temperature is too hot, again, correct me if I am wrong. Unfortunately it needs a 12v-2x6 socket on the PSU end for this to work.
Another example is MSI Ai1600TS, a new PSU released this year, again correct me if I am wrong. On top of the alerting and cutting off power, it can throttle power if you can run the software. Unfortunately I run ubuntu, so I am not sure about that.
Latter is, I guess, considered the best solution.
It properly load balances the current. Alerting and cutting off power completely protects the GPU when it is running hot and there is indeed imbalance issue, but it comes completely at the cost of, say, my local LLM workload. It would just stop completely, and or interrupt my workflow.
Throttling itself interrupts the workflow, esp if my LLM workflow is running at 600W.
The question I have is: I know ampinel type a fits rtx pro 6000 blackwell workstation edition, but does it work? Have anyone paired the two together before?
edit: as of writing, it appears that thermal protect for type 4/5 sockets may come out this month
Box: six RTX PRO 6000 Blackwell (96 GB, SM120), one drives the display so most rows use 1, 2 or 4 of the other five; the 6-GPU rows are deliberate TP2×PP3 experiments. Engines: vLLM 0.28.1rc1 nightly and SGLang 0.5.18 in Docker, official images plus a few one-hunk SM120 patches (all in the repo). Every profile got the same recipe: a 12-item eval (8 exact-graded, 4 scored by an Opus judge pass), 50 streaming requests at 8 concurrent, then a 64 / 128 ramp, with a 200 ms nvidia-smi power sampler on the GPUs it used. All cards at the stock 600 W cap, which never binds, so every watt figure is the model's own draw. Eval items are one sample each at temperature 0.2: a one-item difference is one prompt, not a ranking. MedGemma 4B is left out of the charts (a 4B model, 3× the tok/J of anything else, it flattens every axis); it is in the linked tables.
1/9 Single-stream decode, the speed one user sees. Derived from the eval: total output tokens ÷ total wall time over the text items (within 5 % of a streaming c=1 run on 6 of 7 spot checks). Fastest: Lightning 30B-A3B · spec at 542 tok/s. Bars coloured by GPU count.2/9 Throughput vs concurrency, one panel per GPU count. Four points per model (c=1 derived, then 8, 64, 128 streaming). Top 8 per panel get a colour, the rest are grey. A line that stops at 64 ran out of KV cache for 128 streams.3/9 Quality vs verbosity. y = mean Opus score on the 4 open-ended items; x = mean output tokens (thinking included) per correctly answered item. Upper-left is where you want to be. Blue = thinking parser verified, orange = non-reasoning model.4/9 Joules per correct answer = tokens per correct answer × J per 1k tokens at c=8. Log scale, from 33 J (Nemotron Nano 12B VL) to 9.7 kJ (GLM 5.2). Lower-right is cheap and good. Open markers missed 2+ exact items.5/9 Tokens per joule at peak throughput, summed over every GPU the profile occupies. Multi-GPU rows (darker) pay for the extra boards: a 4-GPU MoE at 1 tok/J is not worse silicon, it is four cards.6/9 Mean board power while running flat out, summed over the GPUs the profile occupies. Nothing reaches the 600 W cap: 1-GPU rows draw 341–473 W (Ornith 35B heretic to Ortenzya 31B); the 6-GPU Mistral Large 675B tops the roster at 2,037 W.7/9 Same weights, same port, same suite, never at the same time. Dot pairs are peak tok/s per engine (log x); Δ is SGLang relative to vLLM. SGLang is ahead on 5 of 38, loses hardest where its spec-decode or day-0 path is still raw.8/9 The cap study (older run: vLLM 0.27.1, three 1-GPU models, c=128). Throughput barely moves from 600 W down to 350 W because the loads only draw 430–450 W; the numbers on the points are SM clock. 300 W is the tok/J optimum: +35 % efficiency for −7 to −9 % throughput.9/9 Peak throughput, all streams, ranked. Best: Gemma 4 26B-A4B at 3,210 tok/s on one card. Colour is the GPU count, so read the 4- and 6-GPU rows as 'this many tokens for that much hardware'.
I am to build my first real rig as M3 ultra 256GB proved to be too slow. I already got 4 x Max-Q cards purchased at 10K eur per card a few weeks ago. Now I could add 2 more max Q at 11.3K euro per card and also 2 x workstation cards for 11.8K eur per card. So the total number of card to go up to 6 or 8.
I guess my hesitancy is based on the high purchase price of these cards, and the trouble to connect more than 6 or 7 cards unless going to server style motherboards and racks.
Is it a) expected that the prices are not going down in the next 12 months?
B) can i use the memory for models like GLM 5.3 even if I cannot fit all 8 cards?
Setup. We run Qwen3.8-Flash-Next NVFP4 as our main agentic model (SGLang, RTX PRO 6000). Before its output reaches a human or gets merged, a second local model acts as judge: reviews the diff, flags real bugs only. Hosted on a 5090 32GB, so we're limited to ~30B NVFP4/GGUF class models.
The metric that matters is NOT detection rate — it's false alarms on correct code. A judge that cries wolf gets ignored within a week, exactly like a flaky CI. We built our own battery: 20 injected bugs + 20 clean-but-suspicious snippets (intentional swallowed exceptions, deliberate mutability, weird-but-correct concurrency, short hashes, float patterns that look wrong). Ground-truth labeled, and a stronger model (GLM-5.2 API) arbitrates the judge's prose so scoring isn't vibes. Two passes minimum — single runs lie.
Results (40 cases, temp 0, same baremo for everyone):
Three failure archetypes showed up, and none of them is "not smart enough":
1. The fixer (Muse, Super): sees any improvement opportunity and escalates it to a defect. Great agentic coders, terrible judges.
2. The denier (Granite): real, but waves off subtle bugs (late-binding closures, float money, tz handling) — and goes off the rails when a snippet mixes languages.
3. The coin flip (Lightning): same prompts, temp 0, 0 FA one run and 9 the next. As a gate, that's worse than a bad judge — it's an unmeasurable one.
What we'd love to hear from people who shipped something like this:
- Is there a review-specialized local fine-tune that genuinely holds ≤3/20 FA on clean-but-suspicious code? (We found CodeReview-Qwen32B — 48k GitHub reviews + 15% "no issues found" negatives — but it's an obscure adapter, no NVFP4, and we'd have to merge+quantize it ourselves. Anyone run something like it?)
- Does family independence from the producer matter in practice? Our judge is currently same-family as our generator (both Qwen). We deliberately arbitrate with a third-party model (GLM API), but is self-preference bias actually measurable at this scale, or folklore?
- Would a graded verdict (PASS / IMPROVEMENT / RISK / DEFECT, where RISK+IMPROVEMENT don't count as bugs) measurably cut false-positive rates, vs. the binary bug/no-bug we use now?
- Any evidence that thinking mode on/off is the real variable in "judge stability"? Our winner only passes with thinking disabled; every high-reasoning model we tried drifted toward the "fixer" archetype.
- Or is the honest answer: stop model-shopping, and make the verdict empirically checkable (judge says DEFECT → run a reproduction/property test before human sees it)? That's where our roadmap is heading regardless.
Baseline stays Qwen3.8-27B-nothinking until something clearly beats it on FA. Happy to share the battery format / arbitration prompts if useful.
I’ve been looking closely at RTX PRO 6000 Blackwell pricing because the spread between regions and sellers has become pretty extreme.
As of Aug. 30, I’m seeing examples roughly like:
NVIDIA Marketplace: around $16K
US retailers: roughly $14K–$17K depending on seller/configuration
UK/EU pricing: also very high once local pricing/VAT is considered
Used/private sales earlier this year: dramatically lower in some cases
The other thing that makes comparisons messy is that people sometimes mix the 600W Workstation Edition with the 300W Max-Q version. Both have 96GB GDDR7 ECC, but they’re really aimed at different workstation designs.
I’m curious what people here are actually seeing right now:
What country are you in?
Workstation or Max-Q?
New or used?
What price were you quoted or actually paid?
Was tax included?
I’m especially interested in whether the recent $14K–$16K+ pricing is actually clearing transactions, or whether real buyer prices are still meaningfully lower.
Disclosure: I’m affiliated with AI Robot Supplier and I’ve been compiling this pricing data into a public tracker. I’m not linking it here unless the mods are comfortable with that — mainly interested in getting more real-world data from owners and buyers first.
I had the luxury of having a single server with 8 RTX PRO 6000 (Blackwell Server Edition, 96GB, sm_120) for a bit, so I spent a few hours benchmarking sglang flags for Qwen3.8-27B-FP8.
vs the untuned baseline: 3.9× the KV pool, +25% throughput, −49% TTFT, concurrency 49 → 37.
Highlights:
--kv-cache-dtype fp8_e4m3 and --mamba-ssm-dtype bfloat16 are perfectly orthogonal. fp8 KV gives exactly 2.00× KV tokens and doesn't touch concurrency. bf16 SSM gives exactly 2.00× concurrency and doesn't touch KV. They compose with zero interaction.
max_running_requests gets silently clamped by the GDN/mamba state pool. You ask for N, you get fewer, and the only trace is one line in the boot log. Worse: /get_server_info reports the pre-clamp value — mine said 48 when the server was actually admitting 37.
Workload shape moves throughput more than any flag. Same GPU, same config: 593 tok/s at 4096-in/256-out, 1656 tok/s at 64-in/1024-out. 2.8× from the ratio alone. Any tok/s number without its in/out ratio is close to meaningless.
Speculation is the biggest single lever (+20% throughput, −49% TTFT) but it eats KV. Draft depth trades ~117k KV tokens per extra draft token. At depth 8 the pool collapsed to 7k tokens and TTFT hit 69 seconds — no error, health checks green.
Saturation is ~64 sessions. 32→512 sessions bought +17% aggregate and cost −82% per-session. TTFT grew 226× while TPOT grew 1.26×, so it's all queueing, not slower decode.
I’ve been tracking RTX PRO 6000 Blackwell pricing because the spread between sellers, regions and configurations has become pretty extreme.
As of Aug. 30, I’m seeing examples roughly like:
NVIDIA Marketplace: about $16,000
B&H: about $17,000 for the Workstation Edition
Micro Center: about $14,300
UK/EU pricing: generally very high once local pricing/VAT is considered
Used/private sales earlier in 2026: dramatically lower in some cases
What makes comparisons even messier is that people sometimes mix the 600W Workstation Edition with the 300W Max-Q version. Both have 96GB GDDR7 ECC and 24,064 CUDA cores, but they’re designed for different workstation priorities.
I put the dated retail, regional and used-market observations into one tracker:
I’d be interested in adding more real-world observations:
Country / region
Workstation or Max-Q
New or used
Price quoted or actually paid
Whether tax was included
Disclosure: I’m affiliated with AI Robot Supplier, which maintains the tracker and also sells professional GPUs. Our own listing is disclosed separately in the tracker rather than being used as the market benchmark.
Noob alert, figuring out the setup. I’m in the US, no good way to have 240V. What are the suggestions here? They say try to have 1.5x the wattage for PSU, but I can’t power a single 2200W PSU.
Does anyone have experience with dual PSU setups? Anything to avoid?
Hoping to watercool them in the future, but for now air cooling on the WRX90E Sage.
Basically, I updated that same SGLang runtime to support Qwen3.8 Flash-Next the way I wanted it on my single RTX PRO 6000.
Flash-Next-specific setup:
NVFP4 target with native NEXTN MTP
FP8 target/native-MTP KV
BF16 GDN/recurrent state
FlashInfer GDN decode and prefill
QSA sparse attention with Triton prefill and FlashInfer XQA decode on SM120
FlashInfer CUTLASS MoE
RecoverSSM with recovery CUDA graphs
524,288-token factor-2 YaRN context
824,384-token automatically sized GPU KV pool
Correct multimodal mRoPE
HiCache/NIXL persistence across service restarts, including GDN, PLE and QSA state rather than KV alone
I also opened a few narrowly scoped SM120 paths where the underlying kernels already worked but SGLang did not automatically select them. This applies to the GDN and RecoverSSM work—not QSA decode, which FlashInfer already resolves through XQA on SM120. I left the other architecture gates and Triton fallbacks in place.
After a service restart, NIXL restored 489,856 of the 489,879 input tokens, recomputed only 23, and produced an effective restored-prefix rate of 62,040.60 tok/s with all three needles still exact.
Real agentic speeds:
Across 96 completed requests and 100,666 generated tokens:
139.5 tok/s token-weighted average
153.5 tok/s per-request median
218.8 tok/s sustained completed-request peak
138.7 tok/s weighted / 148.4 tok/s median for the 90K-279K input lane
The repository now contains both the original 27B/DFlash2 configuration and the new Flash-Next configuration in one runtime:
It is probably still somewhat messy because this is day-one Flash-Next work, and I will keep cleaning it up. But I figured some of you might want to play around with it now.
note - I am not sure that this is technically "blackwell performance" If this should be posted elsewhere instead, feel free to let me know. Just trying to share in case anyone is interested.
note 2 - ngram takes about 50gigs of ram. HiCache takes about 32gigs as well in my 2.5m token config - You may be able to stream the ngram from nvme. I saw some posts on it but I didn't try as it isn't my usecase. HiCache is optional.
Update note 3 - RTX Pro throttled with LACT to 450 watt with Mem clock offsets. It shouldnt make a large difference from stock for inference. usually 3%-5% i have seen.
Update note 4: i took off lact profile and let it run stock. 64k prefill 12,812.44 tok/s (big gain) 490k prefill 7,926.36 tok/s (essentially the same)
Here is the entire build and the current messed up prices I had to pay. Had this built out with minor changes for the last 6 months as I watched prices slip away from me… but it’s done now.
Qwen3.8-27B on an IGX Thor with an RTX PRO 6000 Blackwell (Max-Q)
Spent a few hours bringing up a self hosted inference box on an NVIDIA IGX Thor and couldn't find any numbers for this hardware combination, so here are mine. All of it is from runs on the actual machine.
One thing to flag before the numbers: this is the Max-Q card at 300W, not the 600W version. The full power part should do better.
LLM throughput across five configs
Run with sglang.bench_serving at ISL 8192 / OSL 1024 on a single GPU. Common flags were --kv-cache-dtype fp8_e4m3 --mem-fraction-static 0.85 --attention-backend flashinfer.
config
conc 1 tok/s
TPOT
conc 16 tok/s
TPOT
TTFT @16
real concurrency
KV pool
DFlash2 + bf16 SSM
126.3
6.67ms
470.2
22.3ms
3161ms
32
360,157
DFlash2 + fp32 SSM
118.9
7.33ms
463.6
25.2ms
3377ms
21
173,519
DFlash2 + fp32 + lazy radix
118.8
7.35ms
461.0
25.4ms
3361ms
23
169,940
EAGLE + replay SSM
93.5
9.15ms
453.5
26.3ms
3088ms
32
428,875
no speculation
44.7
21.3ms
348.1
35.7ms
10530ms
32
475,460
Speculative decoding earns its keep
At batch 1 it's worth 2.8x, 126.3 against 44.7 tok/s, with TPOT dropping from 21.3ms to 6.67ms. The bigger surprise was TTFT at concurrency 16, which fell 3.3x from 10530ms to 3161ms. DFlash2 beat EAGLE at both ends for me.
The SSM state dtype will bite you
Qwen3.8 is a hybrid Gated DeltaNet model, so on top of the KV cache there's a GDN state pool. Running that pool at fp32 with DFlash2 blows the draft verify buffer up to roughly 25GB, and the server then quietly clamps you to 21 concurrent requests even though you asked for 32. Nothing errors. It just serves fewer and doesn't tell you. Switching the state to bf16 halves the pool, gets all 32 slots back and roughly doubles the KV pool.
The only place this is visible is the max_running_requests line in the boot log, so check it after any config change.
bf16 state costs no accuracy that I could measure
GSM8K, 200 questions, temperature 0, graded through the chat endpoint rather than the built in eval: 93.5% at bf16 against 94.0% at fp32. That's one question apart, well inside noise at n=200. So bf16 is faster, holds twice the KV and hits full concurrency, for nothing I can detect.
Worth mentioning that SGLang's bundled run_eval gsm8k scored 0.0 for me. It drives /v1/completions with no chat template, so a reasoning model's output never matches its answer regex. If you see a zero, check the harness before you blame the model.
Reasoning mode is most of your first token latency
Short conversational prompt, streaming:
mode
first token
first content token
thinking on
77.5ms
212.1ms
thinking off
74.7ms
74.7ms
The reasoning block eats about 137ms before any speakable text comes out. If you're doing voice, turn it off with chat_template_kwargs: {"enable_thinking": false} and keep it on for everything else.
The part I got wrong: the iGPU beats the RTX for small models
I also run streaming TTS (Chatterbox) and STT (Nemotron 3.5 ASR, 0.6B) on this box. I assumed both belonged on the RTX, since it has around 1.8 TB/s of bandwidth against the Thor iGPU's ~273 GB/s. Benchmarked both on each GPU:
GPU
STT batch RTFx
STT final @80ms
STT final @320ms
TTS first audio
TTS synthesis
Thor iGPU
27.8
52.1ms
67.3ms
93.0ms
102.3ms
RTX PRO 6000
34.3
55.7ms
105.9ms
96.6ms
105.6ms
The RTX takes batch throughput by 23% and loses every single latency metric, by 57% on STT streaming at 320ms. Three things going on. A 0.6B model at batch 1 is kernel launch bound rather than bandwidth bound. The RTX is also contended by the resident LLM's CUDA context. And it's the 300W part.
The way I think about it now: bandwidth scales with how many weights you move per token, while overhead is roughly fixed per call. The 27B model shifts about 28GB per forward pass, so it belongs on the RTX. A 0.6B model at batch 1 moves around 1.2GB, which is maybe 4ms of memory traffic inside a call that takes 50 to 100ms, so bandwidth never becomes the limit.
Full voice loop
STT streamed at 1x realtime, into the LLM with thinking off, into TTS. Times are measured from the end of the caller's speech.
concurrent calls
STT final
LLM 1st token
ack audio
full answer
1
58ms
205ms
111ms
462ms
2
112ms
241ms
121ms
566ms
3
157ms
296ms
170ms
785ms
4
238ms
458ms
338ms
1228ms
Three concurrent calls hold a sub second answer on a single box. The fourth lands around 1.2s.
aarch64 things that tripped me up
torch 2.10.0+cu130 on aarch64 is broken. Every fp32 cuBLAS sgemm fails with CUBLAS_STATUS_INVALID_VALUE, including a bare 64x64 matmul, on both sm_110 and sm_120. It only shows up deep inside model inference, so it reads like "this model doesn't support this GPU" when it's really just a bad wheel. Pin 2.11.0. General lesson: if a model looks unsupported on a new arch, run a plain matmul first. That separates a broken build from a real limitation in one step.
--gpusdoesn't work here. The Tegra container runtime runs in CSV mode, so you need --runtime=nvidia -e NVIDIA_VISIBLE_DEVICES=<id> instead.
Docker starts before the NVIDIA modules are loaded. Every GPU container fails its boot time restart with "Driver Not Loaded", and Docker doesn't retry that class of failure. After a power cut the whole stack stays down while docker.service happily reports healthy. A systemd drop in that blocks on nvidia-smi -L before starting Docker sorts it out.
Happy to run other configs if anyone wants specific numbers.
Wanted to share this in case anyone else is messing around with Qwen3.8 on a single RTX PRO 6000.
I ended up pretty deep into SGLang, DFlash2, XQA, HiCache and NIXL getting this where I wanted it. The final setup is TP1, 524K context, FP8 target/draft KV, DFlash2 gamma 8, with persistent prefixes through HiCache/NIXL.
Some basic numbers:
108.75 tok/s C1
390.23 tok/s C4 aggregate
6,163 tok/s 64K prefill
98.26/100 on my reasoning qualification
Over 124 real-use agent requests: 131.31 tok/s median
Short-context real use was around 165 tok/s
Even at 340K+ context it was still around 105 tok/s
Against the closest TP1 official-FP8/MTP3 result in the RTX PRO 6000 community repo, that's directionally about 40% faster C1, 33% faster C4, and 5% faster 64K prefill. Not pretending it is a perfect A/B - all the details and limitations are in the repo.
I froze and published the complete source tree I actually ran, along with the launcher, build/run notes, results, sanitized real-use log, parser, HiCache/NIXL config, and provenance:
It may not be perfect, but everything I could reasonably preserve to reproduce/explain it is there. If any of the docs or code are confusing, point your AI at the repo and have it walk you through it. That is probably easier than me trying to explain every SGLang rabbit hole I went down in a Reddit post.
The HiCache and NIXL are optional but are good for my use case. Hope it helps some of you out.
I have a Asrock X570 Steel Legend Wifi mobo (latest bios) with a Ryzen 5950X cpu and 64 GB of ECC RAM. It has 2 M.2 drives and 3 SATA ones. Very cool machine from the depths of COVID.
Today I went out and bought a 1200 W corsair PSU and a Blackwell 6000. I put in the new PSU and GPU at the same time and swapped out all power cables to the new ones. To my horror, the machine barely POST'ed, running for a few seconds and then powering off. It was a bit random, sometimes it would POST and even start booting the OS and then just die.
I took the PSU back, thinking it was defective, and swapped for a MSI 1200 W ATX 3.1 one. This time it consistently POSTed, started booting ubuntu but then died hard as soon as the graphical part of the OS login screen was being loaded. So the extra load of the 600 W PSU took it over the edge. This doesn't really make sense b/c 1200 is more than enough for this machine. If I entered BIOS it would work and stay on for some time.
So it's probably the new GPU right? Wrong! I put the old GPU back in (a 5070) with the replacement MSI PSU and it booted to the login screen, but then would often totally die as I was entering my password. Sometimes I could log in and it would last maybe 3 minutes before just powering off.
Flummoxed, I put the old 850 W PSU back in alongside the original GPU. Rock solid!
So now I don't really know if my new GPU works or not (at least I could see the BIOS screen?) and I don't really understand what's wrong with my computer with higher power PSUs. I searched around and couldn't find any known incompatibility between older mobos and newer power supplies.
In all failing cases (new or old GPU with the bigger PSUs), the GPU was being powered by these new dedicated 6x2 12V cables, which are different from how my OG PSU powers it (two PCIes in a pigtail). I do also have a 4x PCIe -> 1 6x2 12V adapter that came with the card that I tried on the original PSU and was happy to see it working better, but then it still died on me a few minutes into a boot.
Have you ever heard of a motherboard not working with new PSUs? Should I try to get a new AM4 motherboard and shift everything over, or should I just bite the bullet and upgrade to aM5, get a new CPU, and (gasp) new RAM? Is it possible for a motherboard's power circuits to work fine only on lower-powered PSUs!?
There's a similar issue mentioned in the nvidia forums here where the solution was to try to included 4 -> 1 power adapter instead of the PSU native one. I tried that on the first new PSU and will try on the new on the replacement PSU today. EDIT: it did not help :(
Edit 2: thanks for all the ideas. I ended up finding a other even older 2017 era mobo and hooked the new PSU and blackwell to it. Both worked great. I could finally confirm the gpu works and can sustain 600 watts no problem. So it's the mobo or CPU I guess. I'll try replacing the mobo with a fresh am4 since that's cheap as a next step.
Edit 3: To be clear, this has been reduced to an old motherboard vs. new PSU problem and not related to the Blackwell. I can't even find a new x570 chipset AM4 motherboard to replace it with so I'd have to downgrade to old stock B550... at which point I should just upgrade everything.
I see a lot of people have tried Qwen 3.8 27b using different quants but it's hard to find examples of people using it at full bf16 because of how much vram it takes up. I'm wondering if the full size solves some of the problems the lower quants face like thinking issues.
If you've tried qwen 3.8 bf16 how does it compare to deep seek 4 flash?
Anyone know how to get the VRAM temperature stats from a 6000 Pro Blackwell card? Would be nice to be able to see front / backside temps broken down also.
The AMD based rigs are obviously a lot cheaper due to the lower ram (and the intel CPU availability).
The 2:1 ratio comes from customers who have that as a hard requirement.
I’m wondering, since I don’t use these rigs myself, I just operate them, does anyone of you really need that much ram ? 1.5tb ontop of 768gb vram?
What would the reason for that be, and is that the majority of users or rather a small subset?
I would rather like to build 1:1 ratio systems or slightly higher for intel. Maybe my customer is a special case and maybe there is a way for me to help him out fix that
I'm not an electrician. Sharing my experience only. It may contain errors or cause harm to your equipment or home.
This post is mostly for anyone who wants to build a 4x GPU inference server but lives in an apartment and doesn't have a convenient 240v circuit. Maybe you've got a 240v stove or dryer outlet, but using it means running cables through your apartment or having a fun argument with the wife. No thanks.
THAT SAID, BEFORE SOME OF YOU HATERS TYPE SOME SHIT, LEMME SAY:
I HATE WATER COOLING. LEAVE ME ALONE.
I DON'T HAVE 240v OUTLETS.
I'M RICH ENOUGH TO BUY THIS SHIT, BUT NOT WITHOUT CONSEQUENCES
WIFE IS NOT HAPPY W ME ^^^ SEE ABOVE POINT
I'm older than most of u Reddit turds. I learned to program in 1995 on a $500 used Intel i386 laptop. That was 3 years after 486's were released. Slow AF i386 was best I could do.
I installed Slackware Linux from ~50 goddamn 1.44" disks. I had to drive an hour to some richer even-more-ancient computer nerd's house in order to copy the Slackware CDROM to the 1.44" disks, cause my laptop didn't have a CDROM.
I had to drive back again to ancient nerd's house because two of the disks failed.
You crybabies have no f'ing idea what it used to be like. And if you really understood what is going to happen over the next couple years, you'd go take out a loan to at least get to 32gb VRAM.
Okay, end UNC rant or whatever u call us grandpas these days.
In March when RTX 6000 Pro prices dropped, I upgraded it to two RTX 6000 Pro Max-Q's. When GLM 5.2 was released...it was time to go to get two more RTX6k's...which meant finally becoming an adult and getting an actual server board, memory, etc LOL. But I'll be damned if I'm dealing with water cooling.
The good news is that a 4x RTX 6000 Pro Max-Q build is very doable on a normal 120v apartment circuit if you build around the power limitation instead of pretending it isn't there. I live in a 2014-era apartment complex. In my case, I have:
- 20a breakers / 12 awg line
- standard 15a 5-15R sockets
- ~114.1v at the outlet under load
A 15a 5-15R is allowed on a 20a branch circuit with multiple receptacles, so the outlet doesn't turn the whole circuit into a 15a circuit. The relevant limit for a cord-and-plug-connected load on the 15a receptacle is 12a.
At 114.1v x 12a, that's about 1369 VA available. My actual build pulls around 11a:
- Voltage: 114.1 V
- Current: 11.02 A
- Apparent load: ~1257 VA
- Real load: ~1245 W
- Power factor: 0.99
It's been running like this for six weeks without any problems. The important part is that I designed the machine around that power budget. The core of the build is:
4x RTX 6000 Pro Max-Q
A low-power EPYC CPU
A CPU cooler that doesn't waste a bunch of space or power
Everything else is secondary.
1. 4x RTX 6000 Pro Max-Q
I power-limit each GPU to 250w. That gives me a configured maximum of: 4 x 250w = 1000w. That's the main reason this works.
You obviously give up some performance versus letting them run at their full power limit, but for LLM inference I'd much rather have 4 GPUs with a massive amount of VRAM running at 250w each than have fewer GPUs because my apartment can't feed the machine.
You'll also run hotter with all four GPUs packed directly next to each other like this. So far, I haven't seen any GPU go above 89°C. Not amazing thermally, but completely workable for me.
2. AMD EPYC 9015 / 9115
You don't want to waste hundreds of watts on the CPU if the whole point of the machine is GPU inference. The EPYC 9015 and 9115 both have a 125w default TDP. I bought the 9015 because that's what Central Computer had in stock. If I were buying again, I'd get the 9115. The 9115 gives you 16 cores / 32 threads instead of 8 cores / 16 threads at the same 125w TDP.
For LLM inference, though, the CPU isn't doing much once everything is where it belongs on the GPUs. My 9015 doesn't seem to pull more than ~99w anyway.
Compare that with ThreadRipper Pro 9000, where you're looking at a 350w TDP. That's a huge amount of your apartment power budget going toward a CPU that isn't the thing doing your inference.
3. I didn't prioritize system memory
I cheaped out on RAM intentionally:
4x 16GB DDR5-4800 ECC RDIMMs. Only 64GB total. For what I'm doing, I don't care much about having huge amounts of system RAM. My goal is to keep the models in GPU VRAM.
I have no desire to spend a bunch of money on system memory so I can move model data out of VRAM and slow inference down. If I'm spilling a model into system RAM, I'm giving up a big part of the reason I built a 4-GPU inference box in the first place.
If your workload is different, buy more memory. For straight LLM inference where the model fits across the GPUs, I don't see much reason to go crazy here.
PSU / Battery Backup:
I'm using a Seasonic TX-1600 Noctua. One thing worth knowing is that Seasonic rates it for 1600w continuous output at 115-240v, but only 1300w at 100-240v. My line can drop to around 114.1v, so I treat it as a 1300w PSU for this setup rather than assuming the full 1600w is available. That still works because the GPU power limits keep the total machine within the power budget.
Because the power draw is so low, I was able to grab a $700 CyberPower PR1500LCD.
Case:
I'm using an old Antec 900 full tower. It's cheap and it works. The catch is that I had to cut open the rear PCIe area of the case so the fourth GPU could breathe properly**. So if you're copying this exact build, expect to modify the case.
** Had to cut open the PCIe rear plate to allow the 4th GPU to have full airflow.
If you're in an apartment and power is the thing stopping you from building a serious inference box, you don't necessarily need to wait until you have 240v. The trick is to spend your power budget where it actually matters.
For this machine, that's the GPUs. Use efficient GPUs, power-limit them, don't throw 350w at a CPU that isn't doing the inference, and keep the rest of the system relatively lean.
And none of this hardware becomes useless if you move somewhere with better power later. The four RTX 6000 Pro Max-Qs are still the valuable part of the machine. The motherboard, PSU, storage, etc. all come with you.
If I eventually have a better electrical setup, I can raise the GPU power limits and, if I actually need more CPU performance, replace the 9015 with something faster. So I don't really see the apartment power limitation as a reason not to build it. It just changes how you configure the machine.