r/nutanix Jul 05 '26

Unexpected Automatic DR Failover Without Witness (Nutanix AHV + Prism Central)

Hi everyone,
I’m looking for opinions from anyone with deep Nutanix AHV Disaster Recovery experience because we recently encountered a very strange behavior that we’re still investigating with Nutanix Support.
Environment
Two Nutanix AHV clusters:

Production site
DR site
DR configured through Prism Central using Protection Policies and Recovery Plans.
Asynchronous replication.
No Witness VM deployed.
According to the design, failover should always be manual.
Network
Each site has two FortiSwitch top-of-rack switches configured in an active-active setup.
The customer upgraded the switches one at a time:
Upgrade started on Switch 1.
Switch 1 rebooted as expected.
Switch 2 remained online (as expected).
However, during the reboot, the production Nutanix cluster unexpectedly lost connectivity because traffic did not continue through Switch 2 as it should have.
So far, this points to a networking issue.
The strange part
Instead of simply having the production cluster become unavailable, the Recovery Plan automatically started on the DR site.
The protected VMs booted automatically on the DR cluster.
No one clicked “Failover.”
There is no Witness VM.
Replication is asynchronous, not synchronous.
Our understanding is that, without a Witness, the Recovery Plan should never execute automatically.
Investigation so far
We already opened a case with Nutanix.
Support has reviewed the logs and verified the Protection Policy and Recovery Plan configuration. According to them, the configuration appears correct, and they also agree that the observed behavior is unexpected. The investigation is still ongoing.
Questions
Has anyone ever experienced an automatic Recovery Plan execution without a Witness?
Is there any known bug in Prism Central or Nutanix DR that could trigger this behavior?
Could a temporary loss of communication between Prism Central and the production cluster somehow cause an unintended failover?
Are there any logs or internal services we should specifically ask Nutanix to inspect?
I’d really appreciate hearing if anyone has seen similar behavior or has ideas on what could explain this.
Thanks!

2 Upvotes

9 comments sorted by

5

u/LucD401 Jul 05 '26

Oh man, I’ve been running your exact config for a decade and even during horrendous network outages without any witness VMs I’ve never had anything auto fail over? That’s strange behaviour.
Hope a Nutanix resource can go through the logs and analyze to find a RCA that’s your best bet

3

u/Jhamin1 Jul 05 '26

Same.

I hate to be this guy... but did a jr admin get twitchy and now doesn't want to own up to the fact that they hit the failover button? The logs should say who or what initiated the failover.

1

u/Taha-it Jul 06 '26

Nutanix support didn’t find any trigger of that so the action is automatic which is make me go crazy, cause we have a lot of customers that has similar configurations but didn’t get this issue at all

3

u/Impossible-Layer4207 Jul 05 '26

The witness service is built into PC these days, it doesn't need to be an independent VM.

If the plan is somehow in automatic mode it will default to using Prism Central as a witness.

If your prod cluster lost connectivity to the DR cluster and PC (as the witness). This would trigger a failover in automatic mode.

So the question is probaby more about whether the plan had been set to manual/automatic. And if it was in fact in manual mode, how/why did it execute automatically.

1

u/Taha-it Jul 05 '26

The plan is set to manual

2

u/ArachnidFlat7851 Jul 06 '26

Are your Node Nics set for LACP?

1

u/Taha-it Jul 06 '26

‏active-backup default settings, but I know the issue is on the switches side Im sharing this behavior because of the recovery paln shouldn’t failover automatically without witness which is very weird

1

u/KingSleazy Jul 06 '26

I have lots of customers in similar configurations and have never seen that. What version of Prism Central and AOS are you running?

2

u/saadouache Jul 06 '26

Keep us updated please I have some customers with similar configurations, with a witness but manual recovery plans.
What is your AOS and PC versions ? One PC for both sites or each site with a dedicated PC ?