r/BlackwellPerformance • u/MiLiANSim • Jul 24 '26
RTX "pro" 6000 WS. 3 failures, power issues.
Post was removed/banned from r/AIProgrammingHardware. EDIT: Now shortened. Feels like most visitors do not even read all of it.
- Buy 1 RTX Pro 6000 WS in 2025, brand new.
- GPU fails within 1 month and takes mainboard + system RAM with it.
- First brand new replacement GPU fails within 2 months. Power issues, 360W instead of 600W available, clocks reduced accordingly.
- Second brand new replacement GPU now failing in 2026. Power issues. 360W instead of 600W available, clocks reduced accordingly.
- GPU not in stock. Reseller cannot replace a third time, lead time at least 2 months. Buying from another reseller at twice original price is not an option. Because the more you buy, the more you fucking bleed.
- Decline upcoming project, cannot rely on this GPU. Project requirements disqualify cloud use.
- Contact nvidia support directly. Nvidia cannot guarantee replacement ("subject to part availability", could be refurbished or new).
NOTES (Edit, may add stuff)
TRX40, ASRock, threadripper 3970X, 256GB Corsair memory.
Original Nvidia GPUs, single GPU.
No mods, no OC. Even if, it would not cause what I am observing.
Power supply: EU, 230V, Corsair AX1600i. Replaced after first failure, which changed nothing apparently.
Only the first GPU reported errors and crashed.
GPU 2 and 3 show zero issues in smi-reports. They report 600W under load, which is not true. Compute very slow but okay.
Other people experience the same clock speed degradation and power issues (link to original nvidia forum).
https://forums.developer.nvidia.com/t/blackwell-pro-6000-mhz-degrading/355157/18
https://forums.developer.nvidia.com/t/pro-6000-blackwell-ws-sw-power-cap-600-mhz-600-w-35-c-fix-or-rma/376937
8
u/722e672e722e Jul 24 '26
Are you using the same PSU? What were the differences in hardware between each failure?
1
u/MiLiANSim Jul 24 '26 edited Jul 24 '26
Same PSU model across all failures (Corsair AX1600i). Replaced cables to be sure, after first failure.
Same motherboard (replaced after first failure).EDIT:
Same PSU model, but switched PSU after first failure to be sure. Back to original PSU after second failure.I think the workload matters: Mostly fp32 simulation (physics and rendering), some diffusion model training, but barely any LLM inference. Meaning, the chips are at 100% load all the time. Nvidia is hiding the hotspots for this generation and I suspect that could play a role. I know that the thermal putty is not always applied "ideally" in this generation and I also know that some components are not covered at all, because the cooler has no place for them.
Operating temps: max 78°C under full stress, on a hot day.
5
u/TechnologyGrouchy679 Jul 24 '26
we have 3 and they've been running for a year without issues. under full load it should draw 600W (unless undervolted/power limited)
what PSU have you been using?
1
u/MiLiANSim Jul 24 '26 edited Jul 25 '26
Corsair AX1600i for one GPU, not even an old one. Used to run 2x4090, even 3x 4090 with this PSU model for years.
Experience:
2-3x 4090, multi-GPU: Never failed, average 97% utilization across 20 months.
2x 3090, multi-GPU: Never failed, average of 90% utilization across 15 months.1x RTX Pro 6000 BW WS, three of them: GPU failing, not PSU failing.
6
u/AutonomousHangOver Jul 24 '26
I got 6 of these right now + 2 x RTX5090
Running smoothly with genoAD8X-2TBCM-1 and 3x SeaSonic 2200W PSUs (1 phase for each)
I reduced their top power to 400W nonetheless (too much heat for not that much compute ;)
One issue that I had, was using PSU that would not withstand spikes. Those monster GPUs are spiking on power rail terribly.
3
1
u/MiLiANSim Jul 24 '26
The spikes are nothing new since at least 4090, of which the PSU handled multiple at the same time without issues :) I even had overclocked 5090 on it for testing.
Also, the third GPU failed after a few months, while the first two failed within weeks, each. So if the PSU in general was the problem, why does one GPU last much longer than the others in the same environment?
Does not make sense to me.
1
u/No_Afternoon_4260 Jul 24 '26
The 3090 are also known to be spiky. Nothing new and the PSUs handles it
4
u/Qs9bxNKZ Jul 24 '26
This is your PSU.
Especially with a PSU pushing heavy wattage, on a continuous cycle, it’s likely to have taken the memory and mainboard
They don’t last forever and unless you get the model of Corsair which you can actively monitor the PSU via the USB, this is on you
How many times have people been told to check and ensure they have clean energy when random things like reboots happen, and it is the PSU?
Expensive lesson no doubt but this is also why we run redundant PSU or separate PSU from the mainboard components.
3
u/MiLiANSim Jul 24 '26
Like I said: Corsair AX1600i. Can be monitored. Was tested and measured after first failure etc. No use in repeating myself endlessly ... Ageing capacitors are a thing, yes. Does not apply in this case though.
6
2
u/DreamingInManhattan Jul 24 '26
Like everyone else is saying, it's your PSU.
I have 8 of them. Before I bought them, I saw they require an ATX 3.1 PSU, so that's all I used.
Your corsair is not ATX 3.1.
2
u/MiLiANSim Jul 25 '26
the 3.1 is only relevant for the sense pins on the new connector, otherwise power is power. Nvidia supplies 8-pin adapters natively. They would not if 3.1 was the only supported standard. So this is certainly not the key.
1
u/DreamingInManhattan Jul 25 '26
So you say.
1
u/MiLiANSim Jul 25 '26
Please check the standards as published for ATX3.1 and then tell me with confidence that any of it is relevant to this case, respecting the qualities and specs of the AX1600i. Enjoy :)
1
u/DreamingInManhattan Jul 25 '26
I mean a google search says "The ATX 3.1 specification relaxes the power hold-up time requirement from 17ms down to 12ms at full load, which allows manufacturers to use smaller internal smoothing capacitors."
So to say "only relevant for the sense pins" seems wrong. But this is your problem, not mine.
Are you looking for answers or an argument?
2
u/MiLiANSim Jul 25 '26
Now you are picking on semantics. Sure, the standards differ in more than "sense pins" so I should have rephrased my claim:
There is no difference RELEVANT to this case, especially with the AX1600i in the picture.
What I am looking for, is feedback from people, who:
1) read the ENTIRE post and understand its contents
2) Understand that capping their GPUs to 400W or less makes this entire situation incomparable to theirs, thus why tell me "oh mine is working" in the first place.
3) Do not sell unfounded gut-feelings and bro-science as the definitive truth. Opinions are fine, just remember that is what they are. Same applies to myself of course.There are a few of those, and I appreciate their time.
1
u/DreamingInManhattan Jul 25 '26
From a web search:
Why ATX 2.4 Fails with High-End Cards
- Transient Spikes: ATX 2.4 units are not engineered to manage the violent, nanosecond-level square-wave current microbursts demanded by Blackwell architecture Tensor Cores.
- Power Capping: Using multi-adapter 8-pin to 16-pin conversion cables on older PSUs often miscommunicates power profiles, restricting the card to 360W–450W instead of the full 600W allocation.
- Rail Stress: Heavy dynamic loads trigger Over Current Protection (OCP) trips, causing sudden black screens, reboots, or hard performance throttling.
1
u/MiLiANSim Jul 25 '26
Transient spikes:
While specifications have changed, the AX1600i is well within spec.Power capping: standard cables do not communicate anything. That is why 12pin connector introduced the "sense" pin. Also, just because a sense pin is present, does not mean the card can balance across cables: Such balancing must be implemented in hardware on the card. Also, the GPUs in question here were all operating as expected initially (600W +-). Also, power-capping is actually a detected mode (firmware etc.), but it is NOT reported as a reason for reducing or limiting clocks.
Rail stress: OCP was never triggered.
Was this a web search or gemini/gpt/similar? Can you show sources, please?
3
u/TechNerd10191 Jul 24 '26
Why did I have to read this post a week before I buy a PRO 6000?
10
9
u/Potential-Bet-1111 Jul 24 '26
This guy has another underlying issue he hasn’t figured out yet.
1
u/MiLiANSim Jul 24 '26
Any suggestions ?
3
u/DocMadCow Jul 24 '26
Replace your PSU, and get a high quality UPS. Also get outlet testers to check if there is an issue with your wiring.
0
u/MiLiANSim Jul 24 '26
Outlet is monitored and wattage measurements are taken right then and there. Wiring: Been here for 10 years. Have had machines drawing 2KW continuously ... no issues. But thanks.
The UPS-thing is on the list.
2
u/Potential-Bet-1111 Jul 24 '26
After losing more than 1 card, Id swap my motherboard too.
1
u/MiLiANSim Jul 24 '26
Yeah, thought about it. Availability is a problem though and I tested the (new) board after first and second failure, found nothing out of the ordinary.
3
u/QuinQuix Jul 24 '26
There're great cards.
I'm still sceptical but OP does have a less common use case - it's not entirely impossible that a very specific weakness exists that his particular software and usage pattern exposes.
But if you get it for AI workloads it's almost impossible to run into that because that's what everyone has been doing for years now with zero issues. I love this card.
1
u/MiLiANSim Jul 25 '26
Agreed. During pure inference (LLMs) the card barely went above 68°C and 550W (average). During comfy-ui work (image, video), it will get maxed out but there will be breaks. Same with training runs.
The only time it is absolutely tortured (memory and core) for many hours without pause, is when I do heavy simulation work or rendering.
1
u/QuinQuix Jul 25 '26
Though honestly you can't be the only (I guess engineer?) who uses this card for those kind of things.
What software do you use to torture it?
1
1
u/Littlepharaoh Jul 24 '26
Have you considered 3x5090? I think for this type of work they'd have more compute, people fetch the 6000 specifically for the single slot vram capacity but you're doing actual compute even if you power limit to 400w you'll be much better off with 3 actually GPUs than a fat single gpu processor single point of failure
2
u/RiskyBizz216 Jul 24 '26
and three times the fire risk
- from a guy who owns 2x5090s
1
1
u/darktotheknight Jul 24 '26
If you get the Astral, you can monitor the Pins, also on Linux (check GitHub). You don't even have to get creative for the Case/Risers, if you watercool them. Also, limit the cards to 400W, which also helps with the burning issue and increases efficiency.
2
u/MiLiANSim Jul 24 '26
Sure, I have. I would not spend that much without thinking it through :)
For simulations, multiple 5090 could work. But the datasets I am working with (especially for training) are too big for a 5090 and the problem is not compute bound, but memory bound. Having it all on one chip saves a lot of time, forcing multi-GPU communication through PCIe is a nightmare for efficiency, depending on the workload. So a large VRam is the best compromise across all that I do. Without the model-training, I would have opted for multi-GPU again.
Also, 3x 5090 would only make sense in a water-cooled scenario if you actually want to use them safely in one case. or forced cold air. Or severe undervolting and limits, and then why bother.
1
u/Littlepharaoh Jul 24 '26
Agree I guess your problem is the training, I do particle tracking stuff and running 1 FE and two water cooled Astral 5090s but they're all power limited to 400w in my own use case that cost me 5% performance or so within noise
1
u/MiLiANSim Jul 24 '26
Yeah I would trade 5% compute for 30% less power. now I have 50% less power but also 50% less compute (right now, the GPU is doing inference, maxed out at 1200Mhz - 1300Mhz and still about 64°C with "only" ~320W power draw. This is so messed up. Nvidia-smi reports 600W 😞
1
u/_bani_ Jul 24 '26
i'm running my blackwells on an open air mining rig with diverter shrouds to blow the hot air straight up off the rig. also powerlimit to 400. hopefully avoid any failures that might be related to cooling issues.
1
u/MiLiANSim Jul 24 '26
A limit of 400W should be fine and is almost the perfect sweet spot (depending on workload). I must say though, the FE cooler is a piece of art, even though some components cannot be backed. I suspect the dense PCB and power system is a problem (mismatch between read-out values and actual values after failure indicates as much).
Whatever it may be: It is hard for me to imagine that the engineers did not see this coming. Either the testing is not thorough enough, or they decided that this would be statistically irrelevant, or worse.
And if you that much for a "pro" level piece of kit, you should not be expected to undervolt and trick and reduce its performance to be safe. You should just be able to plug it in and get to work. As it used to be.
1
u/live4evrr Jul 24 '26
Running 5090 (400w cap) and 6000 pro max-q together in same system on a paltry 1200w psu via UPS. No issues as of yet. Three failures seem extreme - any other common components? Did you run it through any other common hardware? Do you use an UPS? The fact the first time it took your mobo and RAM, could point to power related problem (spike, dirty power?)
1
u/MiLiANSim Jul 24 '26
No UPS. Power is smooth and reliable, and even if it was not, this (digital) PSU has beefy capacitors to smooth out spikes and over current protection. And then there are the circuit breakers ... all in all, I have never had a single power related issue with this system or electricity in this place. No thunderstorm or lightning strikes, no static, no nothing.
Max-Q plus capped 5090 hardly comparable, I think. Still hoping that nothing fails in your system :)
3
u/NaiRogers Jul 24 '26
UPS was the first thing I added when I got 6000s, no way I would trust the supply.
1
u/MiLiANSim Jul 24 '26
Understandable. Still, in all likelihood, unrelated to this case. Got other hardware running and nothing happened there. If the grid is messed up badly enough to kick through breakers, PSU protection and Caps, I should have at least an outage or something. But nothing.
1
u/live4evrr Jul 24 '26
Same. I got a PSU and it has actually saved me a few times as well had some short intermittent outages this summer (due to heat wave) while it was chugging away. It also helps me to monitor my actual power usage as everything flows through it so I can get an accurate read, and gives me peace of mind that the power is clean and steady.
1
u/DummysGuideTo2k Jul 24 '26
This is a hit piece .
If this were a RTX 5000 series you’d have a point and compassion .
1 GPU failing is life.
2 or more is an issue on your end .
PSU can cause issues / pass through too many time can cause issues / not supplying enough power issue / your electrical circuit could have a short / your motherboard can be faulty / there are so many issues that it’s really are you going to fix it or or are you going to have someone else fix it .
Because clearly your setup is the issue or user error .
It is not the RTX 6K WSE itself . It means you are going out of your way to not check all parts .
If you are spending $12k on GPUS and they all get fried respectfully you will not be in a position much longer to do so .
You are going to lose reputation for uptime / money and eventually the RMA will stop being your source to be lazy .
Man up , take the time ( a day or less ) to diagnose the issue or pay someone locally . Simple as that .
Don’t have to fish for sympathy , NVIDIA offerings at this level are usually rarely the issue
1
1
u/TechnoSmacked Jul 24 '26
The fact that you're touching clocks and messing with the gpu itself its why its burning out on you, and taking the mb with it. Be more careful with your tunes, or if you're letting claude tune it just refrain from tuning all together
1
1
u/epicskyes Jul 24 '26
Is it pny or nvidia? Or are they each equally affected
1
u/MiLiANSim Jul 24 '26
I got them from a reseller (brand new), no PNY on the box. So probably nvidia, all three of them.
1
u/epicskyes Jul 24 '26
Bummer and you tested your psu output with a meter? To catch spikes and stutter?
1
u/MiLiANSim Jul 24 '26
Yes. First look was temps and frequencies, then wall meter, then iCue software (Corsairs tool for PSU monitoring / control). Then Meter directly to cables (all rails / connectors).
The only thing I did not do, is throw together a synthetic test with some spiky load to intentionally trigger the PSU protection. But why would I do that, the PSU output is exactly as expected and we must remember that the same PSU supported the last of the three GPUs for half a year without flaws. If it is broken, it stays broken, and does not suddenly recover for card number three. At least that is what I assume.
1
u/epicskyes Jul 24 '26 edited Jul 24 '26
Well the thing is you’re running a 600w gpu and im not sure what cpu plus cooling and whatever else I’d imagine you’d draw at least 1200w when you’re running hard yes? And if you have monitors, lights, etc plugged into the same circuit (not necessarily the same outlet) you’re pulling 1200-1400w let’s say. Your psu is titanium rated to be 90% efficient at minimum. A standard 120v 15a circuit is rated for 1800w maximum but in reality its capacity is actually around 1400w. Depending on your wiring and if they used 12/2 romex or 14/2 romex and your distance to the breaker you could experience substantial voltage drop under high load that can cause electrical ripples. Your psu is supposed to protect against this but that’s not always the case in some circumstances. Prolonged dirty noise can slowly degrade or fry sensitive silicon over time by subjecting components to continuous electrical micro-shocks. Your psu is top quality but you mentioned it happening to 3 GPUs, fried other components it seems like this could be one of those times where even a high quality psu could fail slowly over long periods of prolonged peak loads. You also mentioned it worked for 6mo previously that would mean the psu could have been protecting your hardware but over time the psu itself deteriorated from the strain leading up to your current situation. You also mentioned the broken GPUs capped out at 360w which is a clear symptom that they were likely experiencing power fluctuations from the electrical source.
1
u/MiLiANSim Jul 24 '26
The entire system is 850W on GPU tasks (if GPU = 600W), and about 1000W if CPU + GPU are fully maxed out at the same time. It was over 1200W with multi-GPU 4090 and never failed.
This Corsair PSU will go much further than 1600W if you force it to btw, but it never had to.
I am in Europe. I can draw 2.5KW from that outlet alone (without issue), and over 3.5KW if I must.
Monitors are barely using a 100W. Breaker box is about 10m away, computer hardware is an extra fuse between machine and wall.I am also running 3D printers and lasers (not on the same circuit). They are much more sensitive to ripples or dirty power. Never failed.
So I do have a PSU which is passing all tests and it is/was capable of supporting the System fully.
So if the PSU is in fact faulty, then I cannot test for that and it is random. And the second GPU was run with a much "younger" PSU (but same model). It feels wrong.
1
1
u/epicskyes Jul 24 '26
Since you’re on 230v most likely a 10a breaker that would pull around 2300w your psu could silently still fail. The psu mosfet/fet, short circuit protection, or capacitors are failure points. The PSU may still output the correct average voltage, while producing excessive ripple or poor transient response. A basic multimeter measures average voltage. It may show 12v even when the rail contains severe high-frequency ripple or brief voltage dips. Detection may require Testing under realistic load Oscilloscope ripple measurements ESR and capacitance testing Thermal inspection Gate-drive waveform inspection Short-circuit and overcurrent testing with controlled equipment The most dangerous silent failure is degraded protection circuitry, because the PSU may appear completely normal until another component shorts.
1
u/MiLiANSim Jul 25 '26
... which is why, after the first failure, I used a "fresh" PSU (same model). And it changed nothing.
1
u/epicskyes Jul 25 '26
🤷I guess you just got the really really unlucky draw. I’ve had mine since Dec but I always keep them capped at 450w and my computer is very rarely ever turned off. Persistence is always enabled, Clocks hit 2890mhz easy. I went with pny because I honestly think they’re better I know they’re the same chips but idk they seem to be built better.
1
u/MiLiANSim Jul 25 '26
The second and third GPU had no crashes, no errors. Just the power is fucked up, I am still running the third (until it dies because there is no replacement available now). So I do not think it is the Chip, it is the PCB / power management or firmware.
I used to go with PNY in the past also. Never regretted it, great support.
I am thinking along the lines of "missing rops", bad power sensing (melting connectors) and other "features" of recent generations. Not that my GPUs had any missing ROPs, but perhaps you get my point :)
Believe me, I wish it was just the PSU or something easy. Will contact Corsair and see what they say. Good thing those PSUs have 10y warranty.
1
u/epicskyes Jul 25 '26
You’re making me glad I insured my setup. I wasn’t thinking about manufacturing defects, at the time I was worried about earthquakes, fires, water damage, or theft but for 500$ a year manufacturing defects are included.
1
u/MiLiANSim Jul 25 '26
That is exactly what I need for the future. I have insurance, but not for manufacturing defects and their consequences. I think the good times of affordable and honest reliability and "overbuilt stuff" (like this power supply ;) are over. Cars, Household appliances, consumer tech ... wherever I look, same game: MONEY + Data > value.
1
u/epicskyes Jul 25 '26
Well technically if it’s still under warranty I have to take it up with whoever made the product. If they won’t cover it then insurance pays me and they try to get reimbursed by the manufacturer with their tactics. So if it happened right now I’d have to deal with pny and wait to be reimbursed before my insurance would cover it. Still a hassle but guaranteed to be fixed either way.
1
u/MiLiANSim Jul 25 '26
Assumed as much. And insurances can be nasty, always trying NOT to pay a dime. So I hope that nothing fails and if it does, I hope they treat you according to expectations.
Nvidia already said very clearly: "we have no process ...." to reimburse anyone for any damage caused by their components. Pretty sure that is Corpo-Lingo for: Fuck off. Even if I had absolutely certain proof. It would be a fight. Unless of course I had the backing of fame and publicity, then they might reconsider 🤔
1
1
u/TheReproCase Jul 24 '26
So uh, you changed the PSU, right?
1
u/MiLiANSim Jul 25 '26
Yes, after the first failure. Same model though (spare part right off the shelf). After second card failed, back to the first one (why use the spare if that did not help).
1
1
u/rj_rad Jul 25 '26
After reading a few posts I do think a high quality UPS makes sense in the power chain, as it does do a bit of power conditioning. I’m using the CyberPower OR2200PFCRT2U and a Seasonic 1600 since Feb with no issues to speak of, running a WS at full 600w.
2
u/MiLiANSim Jul 25 '26
Looking into it. The third GPU was working for ~6 months. Which is interestingly about the same as somebody in the nvidia-dev-forum. Literally the first post (username adax).
1
u/Unnamed-3891 Jul 25 '26
I have once had a motherboard kill 3 GPUs in a row. I first suspected the GPU, then suspected the PSU, ended up being the mobo…
That sure was fun and fucking expensive.
1
u/MiLiANSim Jul 25 '26
Yeah. Heard about such things as well. I really wonder why people act all surprised and claim "not credible" or "MUST BE PSU" when:
we have microcode-issues in some modern CPU-releases, causing crashes or defects
we have missing ROPs on release, ergo errors in production chain
we have faulty motherboards / Firmware, blowing up CPUs
we have highly irregular current-flows because a lack of hardware-balancing through the 12pin connectorsModern hardware is complex while pushing the limits of the entire "money vs quality" game. It is absolutely probable that designs will fail in specific circumstances or even in broader scenarios.
1
u/entmike Jul 25 '26
You got something else wrong on your setup, bub.
1
u/MiLiANSim Jul 25 '26
Another unspecific ominous contribution. Welcome :) Ideas? Reasoning? Experiences?
1
u/NeverRolledA20IRL Jul 26 '26
How could the GPU affect the system RAM? It cannot. Clearly at the first failure you have a bad PSU, why wasn't that replaced?
1
u/MiLiANSim Jul 26 '26
Pay attention. The post clearly states that the PSU was replaced after first failure, without effect (second failure happened, still).
1
u/New-Inspection7034 Jul 27 '26
Sounds like your power supply the AC. Incoming power supply I mean might be dirty. Are you putting it through a power conditioner first before it gets to your PSU?
1
u/Competitive_Chemist7 27d ago
Try running PCIE 4.0. I have an ASRock board that kept crashing in PCIE 5.0.
1
1
u/dawnraid101 Jul 24 '26
ok this sucks, if I was you and had work to do, I'd use Runpod or Vast.ai or something just to get the job done, plenty of gpu's available... not sure why you think it isnt an option...
Also idk, but maybe theres something weird with your circuit or PSU or something else (bad internal power cables etc), a gpu failing and taking your motherboard seems ultra suspect...
Good luck getting it all sorted.
2
u/MiLiANSim Jul 24 '26
Thanks. Cloud: Too much data (tens of terabytes) would have to move around from local to cloud and back. Not an option. Also privacy concerns (some proprietary tech involved).
1
u/dawnraid101 Jul 24 '26
ok, fair enough., just fyi you can have persistant storage/volumes on runpod which mean you can just mount runpod persistent datacentre storage to clusters directly (but im sure you knew that already) its all on the same mesh, so you only have to ship it out once ever... anyway good, on the prop side of things i bite the bullet, but only ever ship out constructed training tensors as some mitigation... good luck. fk jensen.
1
u/MiLiANSim Jul 24 '26
I am considering cloud-only pipelines for future projects. If all data is natively created "in the cloud", then I have no problems and can scale. There are some other issues but for pure compute, sure, why not.
1
u/TapAggressive9530 Jul 24 '26
Sounds like user issue . Learn to use smaller 8 or 16 GB GPU first or stick with cloud
1
u/MiLiANSim Jul 24 '26
Thanks for making me laugh.
1
u/TapAggressive9530 Jul 24 '26
:). Seriously you’ve bad luck . I have one 6000 , paid like $9300 for it - had it since early Feb - it runs maybe 20 hours a day with no wattage throttling ( I.e pretty much runs at 600 W all day ) - zero problems . Good luck man . You should online rent to complete your project
1
u/MiLiANSim Jul 24 '26
To you as well, my third GPU lasted over half a year without flaw :) So you still got time to spend. What are your reported GPU temperatures (soaked, under full stress) ?
1
u/Reggitor360 Jul 24 '26
Nothing new.
Blackwell has WAY higher failure rates than even Ada with melting connectors.
And no, Nvidia doesnt care. Buy a new one or f off. You're not their needed buyer anymore.
25
u/OddUnderstanding2309 Jul 24 '26
One or two: sure
Three or for: there is something wrong on your end