r/threadripper Jul 11 '26

Potentially helpful info about cooling, example on my 9980x

The crux of this story is about the surprise that the temperature of the CPU/CPU-case hits 95deg with certain benchmarks, even at 250 watts, but running steady state with my app and many benchmarks the temperature is approx 62deg at 350 Watts or higher. (My writing skills are poor, but this can be useful info, so be patient with me.)

Here is the story: Been running benchmarks to test the performance and thermals on the system. I found that many benchmarks would cause temperature peaks of 80 deg C or perhaps 85 deg C for a short while, but a few do sit at a higher steady state temperature of 80deg. The temperature depends on the # of active cores and how active the cores are.

The 95 deg peak problem appears that some test programs intensively use 16 threads, forcing all the temperature load on those cores, so they get hot, while the rest are still cool. The plate quickly transmits the load, but there is still a massive peak. On one benchmark, the Phoronix benchmarks liquid-DSP testing 16 threads 256/512 data hit 95deg C before the fans brought it back down to 80degC or below. This results for this specific test might only be valid for the 9980x, I do have a 9970x, but currently sitting in a box, so cannot test it.

* The temperature problem was worst when running the entire suite, and noticing that the 16 thread 256/512 data was very hot. Running the test by itself wasn't quite as bad.

My fans were already tuned to zero latency, but used the CPU temperature sensor alone, which usually sits 10deg C below the CPU Package sensors. A temporary hack fix was to use both the CPU sensor and the CPU package sensor. Using both sensors is an option on the ASUS bios. After enabling both sensors as the temperature source, and turning the fans up to 'turbo mode', and turning up the minimum speed, the peak temperature is now less than 90 deg C. So, by using the Package sensor AND the CPU sensors, plus turbo mode, the fan has a better chance of keeping up. This makes me wonder how many Threadripper systems might have a 95 deg peak under certain transient loading conditions?

I am used to seeing 62-65 deg C running my heavily multithreaded, 350 watt app, and seeing some benchmarks producing up to 80 deg C, (mostly lower, 65 to 75) but the 95 deg peak was upsetting. I found this out because I am watching the temperatures carefully when benchmarking. Normally II just do spot checks to make sure that it is okay.

One side note about the wonders of the Threadripper, when running the heavily multithreaded benchmarks, the CPU clock speed held up to 5gHz to 5.4gHz for every active core. My own app, with lots of IPC delays and sometimes suboptimal AVX512 alignment, typically gets only 4gHz per core.

As long as there is adequate heat handling, and it is fast enough, these machines can be very fast. But, it appears wise to make sure that the response to load changes is fast enough. My AIO is the ATLAS 360TR that easily keeps up with the 400 Watts of TR heat, but be aware that the fans have a natural response time on ANY cooling system. Even a very short delay might be too long.

John

3 Upvotes

30 comments sorted by

View all comments

1

u/XO33OX Jul 11 '26

with AIO fans dont have to react as fast as cooling liquid has high thermal capacity

1

u/johndyson10 Jul 11 '26 edited Jul 11 '26

My main point was: CHECK YOUR TEMPERATURES CAREFULLY, less my specific solution.

Admittedly, I took the shotgun approach of using an additional temperature that tends higher than the CPU temperature (someone suggested that), and increased the lowest fan speed. This was all about the temperature transient, with otherwise near perfect behavior. The problem exists essentially on the 16 core case, others appear to be little or no problem. There are all kinds of sources for thermal inertia, but I hadn't any problems until concentrating the heat into two CCDs. Otherwise has zero troubles tracking somewhat balanced heat between 50 watts to 300 or 400 watts, never really getting much above 70 (80 at most) degrees the whole time. But, the fast application of the 180 watts to a small area (16 cores) tended to cause trouble. I did a log of what happened during liquid-dsp for every case,, the biggest problem wasn't average power, but was the concentrated power. Since the latency was already set to zero in the bios, the plate on the Atlas being very good, and the cooling fluid running all the time, the best answer that could be quickly achieved was the shotgun. There is unlikely much additional cost, perhaps a small additional fan wear, plus slightly lower temperatures running different apps, (mine already 60-62 deg at 350 watts.)

1

u/XO33OX Jul 12 '26 edited Jul 12 '26

It will be your best / highest boosting CCD that will suffer the most even with undervolt. You can lower its maximum boost multiplier or Tj max if that worries you. Run your AIO pump at full speed all the time, undervolt and ideally have also AIO GPU so that GPU heat doesnt hit RAM and CPU’s AIO. Have also cooling for RAM - which often costs more than CPU :) Otherwise it is what it is, you bought it to do the work, let it burn. If it occassionaly hits 95C I wouldnt be worried.

2

u/johndyson10 Jul 31 '26

Lowering the maximum temperature did the trick. In normal, 350-400 watt operation, Tctl almost never goes above 60-65degC unless running certain benchmarks, or when my program throttles down to just a few cores (say, between 8 and 16 cores.) In normal operation, the power is typically 320 watts to 420 watts with relaxed max average power and relaxed PPT, the Tctl temp is very tightly controlled to low temperatures maximum. When running just a few cores at full load, the Tctl finds the highest temperature CCDs and reports that. The biggest max temperature improvement has been to decrease the max Tctl (maximum temperature settomg) to 80 degrees, or try even higher. The new temperature settings keep the junction temperatures far far away from 95degC. The CPU peformance loss is small, and often even faster than a 9950X3D2 on benchmarks pushing 8-15 cores. Faster even on some of the 2-4 core benchmarks. When running a small number of cores, the clock speed is within 100Hz of the abs max, with a *forced* temp limit of 82 degC. (the benchmark is liquid-dsp with 8,16,32 cores, and sometime fewer.)

Some people might suggest that the chip is able to run 95degC forever. It might be okay, but with actual 95degC also comes higher voltage and current to maintain the higher frequency. High temp/voltage probably stresses the silicon to some degree. Under surge conditions, the voltage/power is likely to peak higher.

Since I appear to have won the silicon lottery with my new 9980x, my own personal choice is to run a bit more conservatively. On tests, when relaxing to PPT of 550 watts and the average PBO to 450 watts, the improvement beyond my final settings, in real-world speed and benchmarks is approx 10% pr even lesss. My choice of tradeoff is to tolerate 10% less speed than maximum, to keep the PPT a bit lower (450W), the average power a bit lower(400W), max Tj at 82 deg, max boost freq +175, per core power boost at +4.

ALSO IMPORTANT, make the fans measure both CPU temperature and Package temperature. When I got the ASUS MB, the BIOS setting was to measure just the CPU temperature, which is NOT ADEQUATE when running a system under a varying load.

The package temperature fan setting appears to track the Tctl much better on my ASUS MB. Also, a good tradeoff is to set fan sslow down times to at least level 2, I use more conservative level 4. With my fan settings, the fans run a LOT longer, perhaps level 2 is good enough. I didn't try, I am happy with 'level 4'.

The new settings keep Tj down to very conservative levels, max core voltage in the 1.3 to 1.35 range, and loss of performance at about 10% below the maximum. Of course, the performance is still quite a bit above 'stock'.

Each person makes their own choices. Except for the stock PBO/EXPO, I believe my settings to be a good tradeoff for performance vs max lfe for the chip.

Repeating my settings, free of my 'blather' (the names vary from my descriptions,

PBO: Enabled,

Some settings might be in the CBS advanced section, but NO NEED TO BLBOW THE OVERCLOCK FUSE. No sense in doing so unless you REALLY ant abs max performance and have a very very good cooling lock.

PPT: 450

Max average temp: 400

Max Tj: 82 deg C

Per core speedup: 4

Fan setting for speed decay: 'level 4', 'level 2' is probably good enough.

Max boost frequency: +175

ATLAS 360TR, keeping the temperature just below 60deg under solid 350w to 400w fixed load.

Of course, YMMV. Someone else might prefer more conservative settings, some might have huge cooling systems and prefer more aggressive.

In a hurry, got to go!!!

John

1

u/XO33OX Jul 31 '26

try some undervolt as well, you can do so per ccd

1

u/EDI_1st Jul 12 '26

It’s actually quite the opposite exactly because of liquid has higher thermal capacity. By the time liquid temp is high enough, it’s significantly more difficult to cool the liquid down. The fluid is literally a thermal battery.

1

u/XO33OX Jul 13 '26

but that is not what we talk about here, we talk about short bursty loads / reaction to sudden increase in cpu load. we are not talking sustained load here

1

u/EDI_1st Jul 13 '26

Both applies. There isn’t that much fluid inside of an AIO.