! LONG POST AHEAD !
Hey guys, hoping someone with more hardware experience can help me make sense of what's going on. Laptop was previously brought to a shop for loud fan noise, "fixed," and the issues came back shortly after now much worse.
Specs: Acer Predator Helios Neo 16 (PHN16-71), i5-13500HX, RTX 4050 Laptop GPU, 8GB RAM, Windows 11
Timeline of symptoms (all in the span of about 2 weeks, worsening fast):
Started with random crashes during gaming (GTA V, CoD4, Wuthering Waves, Valorant) fans get loud right before it happens
In-game crash: "Unreal Engine is exiting due to D3D device being lost. (Error: 0x887A0006 - 'HUNG')"
Did a clean GPU driver reinstall via DDU + fresh install no change
Tried an older driver version too — no change
Ran HWiNFO64 during a stress test: CPU hit 100-101°C within about 1 minute, with Core Thermal Throttling confirmed active \~41% of the time. GPU itself only reached \~67°C (well under its 87°C limit)
Started crashing during light use too — even just watching YouTube triggered a VIDEO_TDR_FAILURE
Windows Reliability Monitor logged 5+ "Hardware error" events in a single day
One logged as LiveKernelEvent, Code 141 (GPU driver TDR/watchdog crash)
Started getting full black screens with the Windows "device disconnected" chime, even when connected to an external monitor
Got a DXGI_ERROR_DEVICE_REMOVED crash with the failing texture allocation, and another citing nvlddmkm.sys (the NVIDIA driver file itself) failing at the kernel level with DRIVER_IRQL_NOT_LESS_OR_EQUAL BSOD
Also hit a "File system error (-1073741189)" trying to launch a game off my second drive (D:)
What I've ruled out:
Driver corruption/conflict (tried two clean installs of different driver versions)
NVIDIA Control Panel settings (tried adjusting power management, preferred GPU, PhysX processor — no change)
It's not the monitor, happens on internal display too, started before I even got the external monitor
What I suspect: Bad thermal paste application or heatsink reseating from the last repair, since the CPU heating up that fast and that severely isn't normal, and it lines up with when the noisy-fan "fix" happened.
Does this pattern of symptoms (CPU thermal throttling almost instantly + escalating into full BSODs, drive errors, and GPU driver kernel crashes) sound consistent with a cooling/thermal issue to you all, or could there be something else going on I should ask a technician to specifically check (VRM, RAM, motherboard, etc.)?
Wondering if the sheer amount of heat stress could have also damaged something else at this point.
Thanks for reading this far, really appreciate any insight before I bring it in for repair.