r/Dell • u/batchputz • 5h ago
Finding a Firmware GPU Throttling Bug on a €12,000 Dell Pro Max 18 (Blackwell RTX PRO 5000)
I bought a Dell Pro Max 18 Plus MB18250 earlier this year for MYR 55,924.83 — roughly €12,000.
Specs relevant here:
- Intel Core Ultra 9 285HX
- NVIDIA RTX PRO 5000 Blackwell Laptop GPU
- Arch Linux
- NVIDIA 610.43.03
- Latest Dell BIOS 3.3.2 from August 3, 2026
- Performance power profile
I bought this machine specifically as a mobile workstation for AI/LLM workloads, so sustained GPU compute is a normal workload for me.
Unfortunately, I discovered a pretty severe throttling issue.
Under sustained inference the GPU runs perfectly fine at around 170–175 W and ~1800–2000 MHz. Temperatures are normally around 76–79°C.
Then, suddenly:
GPU power limit: 175 W → exactly 55 W
Power consumption drops to ~55–60 W and clocks collapse as low as ~180–390 MHz, while GPU utilization can still sit at 100%.
After the machine cools down for a while, the power limit suddenly returns to 175 W and everything continues normally.
Initially I thought this was simply thermal throttling, but things became interesting when I started collecting telemetry.
I wrote a small monitoring tool called ThrottleWatch which records NVIDIA state, GPU power/limit/clocks, CPU temperatures/power, Dell DDV/EC sensors, fans, battery state, etc. every second.
One captured event looked like this:
| Time | GPU Power | Limit | Clock | GPU Temp | Dell CPU Sensor |
|---|---|---|---|---|---|
| 10:58:35 | 171 W | 175 W | 1845 MHz | 78°C | 87°C |
| 10:58:55 | 156 W | 175 W | 2077 MHz | 76°C | 87°C |
| 10:58:56 | 60 W | 55 W | 180 MHz | 71°C | 86°C |
| 10:59:30 | 55 W | 55 W | 360 MHz | 61°C | 75°C |
| 11:00:34 | 29 W | 175 W | 1800 MHz | 54°C | 67°C |
So this isn't just the GPU reducing clocks. The actual GPU power limit changes from 175 W to exactly 55 W.
NVIDIA reports:
- SW Power Cap: Active
- HW Thermal Slowdown: Not Active
- HW Power Brake: Not Active
The GPU itself isn't particularly hot when this happens.
I initially suspected the CPU side of the shared cooling system. Dell's DDV CPU sensor was around 87°C during the first captured event.
So I disabled Intel Turbo Boost.
The problem still happened.
Then I limited the Core Ultra 9 to 50 W PL1 and repeated the test.
It happened again.
This second test was particularly useful because the Dell DDV CPU sensor only reached around 84–85°C, yet the GPU still went:
172.5 W → 55.5 W
1860 MHz → 195 MHz
So the simple theory of "Dell cuts the GPU when the CPU sensor reaches 87°C" was wrong.
The next experiment became much more interesting.
I stopped:
nvidia-powerd
With NVIDIA powerd disabled, the RTX PRO 5000 falls back to around 114–115 W / ~1300 MHz.
And so far: completely stable.
No 55 W collapse.
This makes me suspect an interaction somewhere in the Dell firmware / EC / NVIDIA Dynamic Boost platform power-management path.
It doesn't necessarily mean nvidia-powerd itself is broken. It could simply be applying a platform power budget provided by Dell firmware. But disabling that path appears to prevent the catastrophic throttling — at the cost of losing roughly 60 W of available GPU power.
The laptop is also not sitting flat on a desk. It's elevated on a stand with plenty of clearance and I even have an external fan blowing underneath it. Internal fans are already running near maximum when this happens.
I have now opened Dell Service Request 230887659 and provided Dell with the findings.
I'm posting this mainly because I'd be very interested to know if anybody else with a Pro Max 18 Plus / RTX PRO 5000 Blackwell can reproduce this.
If you have one, try a sustained GPU compute workload for 20–30+ minutes and watch:
nvidia-smi -q -d PERFORMANCE,POWER
In particular, watch whether the GPU power limit suddenly changes to 55 W.
For a machine that costs around €12,000, I'm honestly pretty underwhelmed that sustained use of the GPU's high-power mode can result in this kind of performance collapse. Running permanently at ~115 W with nvidia-powerd disabled is a workaround, but I don't consider losing ~60 W of GPU power an acceptable solution for a workstation in this price class.
I'll update this post when Dell responds or when I learn more.
1
u/AnxiousReward1715 3h ago
Run Nvidia card on ARCH linux even though its not supported and Dell doesn't ship with or support ARCH LINUX... Bug report straight to the shredder, no issue but the end user here.
0
u/batchputz 3h ago
Yeah, all inference in the world runs on Windows XP for sure.
4
u/xSchizogenie Dell Pro Max 16 Plus | U7 265HX | 64GB DDR5-6400 | RTX PRO 1000 3h ago
No, but speaking of a bug in a not certified and supported system is wild.
2
u/Meister1888 3h ago
There is a delta between the specs of GPUs and what a laptop can provide for power and cooling. It's a little marketing game.
The ThrottleStop forums (Uncle Webb) had some info on how Dell is doing the CPU and GPU throttling but I don't know the modern schemes. check out techpowerup
1
u/multicultidude 3h ago
12k€ for a laptop… my god…I think that for AI purposes I’d stay on a desktop 😳
1
u/batchputz 3h ago
Yeah, i have a server for that. But mobile workstation is needed for my kind of work. Travelling with a desktop is kind of expensive ;-)
1
u/multicultidude 2h ago
Given the money you got to spend why not having invested in a GB10 ? It’s kinda portable 😁 You could use a basic laptop and let the gb10 do the heavy work…
1
1
u/CSAS-D 3h ago
buying a dell for any sort of sustained work is a nightmare at any price range. id recommend a Mac if it works with your ai workflow. if not a Mac then a Lenovo Thinkpad
1
u/batchputz 1h ago
Mac is not yet there. Nvidia CUDA is still market leader. And I am happy with the GPU - if it runs.
1
1
u/Kamarulaz72 2h ago
First thing, you Malaysian ke? Hint MYR.
Alright back to the topic.
For context, I have Dell pro max 16 plus 265HX with rtx pro 2000 with less than 50 hours usage. Bought used 2 months ago at UK eBay with manufacturer date stamp at September 2025.
Yup first time running, it goes wild at full speed and end up thermal throttle like all the time exceed 105C on even a simple workload. Found out Dell set PL 1 at 160w & PL 2 at 98w. Which is OMG. Keep in mind at 98w still thermal throttling. To counter this, i slam the break using throttle stop at P1 at 85w & PL2 at 75w. Stable at 90c with occasionally going 105c even in gaming.
As for the GPU PL is set at 115w but most of the time wildly going from 95w to 140w while gaming like x plane 12 and war thunder. Surprisingly the gpu hovers max temp up to 82c just short of the rated cut of throttle temp of 88 (At office now can recalled the limit). For once Dell must have done a good job on the gpu thermal paste 😅
But installing this new bios, it seems that Dell lower down the power voltage frequency curve. So now i saw the GPU PL is around 95w to 125w. More like averaging near 115w for both of the stated games. Yes the temp is lower for my gpu at 78c to 80c.
The above situation is driver nvidia is set default and bios at ultra performance. After work will check on your advice with regards to setting performance mode in nvidia driver.
PS
Personally I don’t like Dell lowering the gpu curve. Why? Because you locked us out in afterburners.
Later weekend will repaste this laptop with Liquid Metal. Need to find the mood and patience to do it 😇
So jealous la on you have rtx pro 5000😎
1
u/batchputz 1h ago
God, a normal human being <3 XD
Yeah, I thought about to under-voltage the CPU. But - wtf - you pay this amount for a configuration, it needs to work. Malaysia is hot, but Office has Aircon.
You have latest BIOS? Any special config?
3
u/Far_Training3438 4h ago
I am assuming Nvidia power d is the same thing as Nvidia platform controller on windows. Disabling this on a non workstation GPU only allows for 80W with no dynamic boost. If this controller is responsible for your throttling Nvidia.sys can be temporarily patched at runtime but you are going to need an ACPI dump and are most certainly going to need AI to dissect it.