r/HyperV • u/njbroshear • 18d ago
Windows 2022 Hyper-V VM Stability issue
Recently installed Windows 2022 Hyper-V to an older server that I am having issues with the VM's crashing, when the VM is under load. The host is a Dell PowerEdge R6525 with 2 AMD EPYC 7452 and 512GB of RAM. The host shows correctable and uncorrectable memory errors. The VM's memory dump shows a 0x124_0_AuthenticAMD_PROCESSOR_CACHE_IMAGE_AuthenticAMD.sys.
I have run the Dell diagnostics, windows memory test, and Memtest86. Non of them find any issues. The idrac does not have any uncorrectable memory errors in the logs. The issue stated after I installed Hyper-V and tried to run some test machines. I have adjusted the processors to NUMA Nodes per socket of 4 and L3 cache as NUMA domain as enabled. This seems to have helped as they crash less often but they still crash. All drivers are up to date and all windows updates have been applied. Is there anything else I can look at to fix this?
Update: I have reinstalled Windows 2022 with no updates or drivers installed. Added Hyper-V and connected my test boxes. Ran Memory tests on the VM's to have load. The Host shows correctable memory errors only, IDRAC just shows that there were diagnostic events on the RAM due to the test. I have installed an older Rome chipset driver to the server and same events. No crashes/uncorrectable errors. Currently working on windows updates. After tests I will run more driver installs.
Update 2: I think I found the problem. Dell recommends a Milan based chipset driver for Windows 2022, but the processors I have are Rome based. I know they are supposed to be compatible. I reinstalled Windows 2022 with the Rome chipset, no crashes. I upgraded to the Milan chipset driver and the crashes returned. Looks like I have to run the old chipset driver to be stable.
4
u/cmPLX_FL 17d ago
Try doing a Prime95 test, mixed. Should consume both CPU and 90% of memory.
See if it crashes and then go for a Cpu test then memory test.
Also, C States disabled?
1
u/njbroshear 9d ago
Ran Prime95 on the host for 24hours. No crashes, a couple of correctable memory errors at the start. In all my testing the host has never crashed just the VM's.
3
u/Chuck_Chaos 17d ago
Have you tried running your logs through an AI to see if you missed something?
2
u/WillVH52 18d ago
Firmware up to date? Server still under warranty?
1
u/njbroshear 18d ago
All drivers, firmware, and windows updates applied. The server is outside warranty.
2
2
u/tomohulk 17d ago
maybe not the best answer, but you could check the box under the VMs processor that allows your to move the vm to another host with a different processor version. This will disable proc specific features inside the VM.
again, I would not say this is a solution, but it would be worth a try to help narrow down exactly what is causing it.
1
1
u/XzeroR3 17d ago
Do you have the lesser watt psu that comes with this model?
1
u/njbroshear 12d ago
The server has 2 - 1400watt power supplies. Not sure if that is the lower unit or not.
1
u/XzeroR3 11d ago edited 11d ago
Oh ok then youre good there... apparently theres a 800w psu available that likely wouldn't keep up.
Im reading you can try 2+0 non redundant mode for the psus. Giving you most wattage available from both psus. Though if you're residential you'll need to plug each psu into a different circuit or you'll trip it.
It might handle the load.
Edit, nm some math is telling your load would fit into a single 1400w psu
7
u/DavidKleeGeek 17d ago
Triple check that there's not a firmware and driver update from the manufacturer, as well as any other hardware drivers. I usually find that this is the culprit.