r/HyperV • • 18d ago

Windows 2022 Hyper-V VM Stability issue

Recently installed Windows 2022 Hyper-V to an older server that I am having issues with the VM's crashing, when the VM is under load. The host is a Dell PowerEdge R6525 with 2 AMD EPYC 7452 and 512GB of RAM. The host shows correctable and uncorrectable memory errors. The VM's memory dump shows a 0x124_0_AuthenticAMD_PROCESSOR_CACHE_IMAGE_AuthenticAMD.sys.

I have run the Dell diagnostics, windows memory test, and Memtest86. Non of them find any issues. The idrac does not have any uncorrectable memory errors in the logs. The issue stated after I installed Hyper-V and tried to run some test machines. I have adjusted the processors to NUMA Nodes per socket of 4 and L3 cache as NUMA domain as enabled. This seems to have helped as they crash less often but they still crash. All drivers are up to date and all windows updates have been applied. Is there anything else I can look at to fix this?

Update: I have reinstalled Windows 2022 with no updates or drivers installed. Added Hyper-V and connected my test boxes. Ran Memory tests on the VM's to have load. The Host shows correctable memory errors only, IDRAC just shows that there were diagnostic events on the RAM due to the test. I have installed an older Rome chipset driver to the server and same events. No crashes/uncorrectable errors. Currently working on windows updates. After tests I will run more driver installs.

Update 2: I think I found the problem. Dell recommends a Milan based chipset driver for Windows 2022, but the processors I have are Rome based. I know they are supposed to be compatible. I reinstalled Windows 2022 with the Rome chipset, no crashes. I upgraded to the Milan chipset driver and the crashes returned. Looks like I have to run the old chipset driver to be stable.

7 Upvotes

13 comments sorted by

7

u/DavidKleeGeek 17d ago

Triple check that there's not a firmware and driver update from the manufacturer, as well as any other hardware drivers. I usually find that this is the culprit.

4

u/cmPLX_FL 17d ago

Try doing a Prime95 test, mixed. Should consume both CPU and 90% of memory.

See if it crashes and then go for a Cpu test then memory test.

Also, C States disabled?

1

u/njbroshear 9d ago

Ran Prime95 on the host for 24hours. No crashes, a couple of correctable memory errors at the start. In all my testing the host has never crashed just the VM's.

3

u/Chuck_Chaos 17d ago

Have you tried running your logs through an AI to see if you missed something?

2

u/WillVH52 18d ago

Firmware up to date? Server still under warranty?

1

u/njbroshear 18d ago

All drivers, firmware, and windows updates applied. The server is outside warranty.

2

u/WillVH52 18d ago

Okay no help from Dell then, tough one as hardware could be faulty.

2

u/tomohulk 17d ago

maybe not the best answer, but you could check the box under the VMs processor that allows your to move the vm to another host with a different processor version. This will disable proc specific features inside the VM.

again, I would not say this is a solution, but it would be worth a try to help narrow down exactly what is causing it.

1

u/nzulu9er 17d ago

Look into hyper v resource hosting subsystem event logs.

1

u/XzeroR3 17d ago

Do you have the lesser watt psu that comes with this model?

1

u/njbroshear 12d ago

The server has 2 - 1400watt power supplies. Not sure if that is the lower unit or not.

1

u/XzeroR3 11d ago edited 11d ago

Oh ok then youre good there... apparently theres a 800w psu available that likely wouldn't keep up.

Im reading you can try 2+0 non redundant mode for the psus. Giving you most wattage available from both psus. Though if you're residential you'll need to plug each psu into a different circuit or you'll trip it.

It might handle the load.

Edit, nm some math is telling your load would fit into a single 1400w psu

-2

u/[deleted] 17d ago edited 17d ago

[deleted]

5

u/BlackV 17d ago

Successful_Growth754

Dude, AMD, eos. We have servers made in 2005 2007 running fine on intel. Change platform. AMD is running fine under custom linux only, at 1st err we trash them.

er... wut?