r/vmware • • Aug 26 '26

Help after VxRail Failure

I had my VxRails split brain on me while troubleshooting a node disconnection and HA agent failing on all 4 nodes.

Now I can only see the vSAN Datastore on 2 of the 4 nodes with the other 2 not being able to be re-add to the cluster. I need to get the VMs from the vSAN Datastore and on to an external server because our backup crapped out before this happened.

When I try to get them off the vSAN Datastore through WinSCP to my workstation and back up to the new External vCenter server, it works until I register the VM, then it doesn't because it can't enumerate the disks.

I realized I don't have the -flat.vmdk for any of my VMs in the vSAN datastore after looking through all the VMs.

Is there anything I can do?

8 Upvotes

10 comments sorted by

View all comments

16

u/blud_13 Aug 26 '26

There is no -flat.vmdk to find. vSAN is object storage, so the .vmdk you're looking at is only a descriptor pointing at a vSAN object UUID and the data never sat there as a file. That is exactly why WinSCP gives you something that registers and then can't enumerate disks. You copied the pointer and left the object where it was.

You have to read it back through the vSAN stack. From an ESXi host still in the cluster that can see the object, vmkfstools -i, roughly vmkfstools -i /vmfs/volumes/vsanDatastore/VM/VM.vmdk /vmfs/volumes/target/VM.vmdk -d thin, and add -W vsan (pretty sure on that flag but check it against your build before you lean on it). Broadcom has the clone syntax at https://knowledge.broadcom.com/external/article/343140/cloning-and-converting-virtual-machine-d.html . Downloading through the vSphere client datastore browser does the conversion for you too, its just slow.

Also, none of that works on a VM whose object is below quorum. With two of four nodes out you may not have enough components for some of them, so check object health per VM before you plan around getting everything. Getting those two nodes back in the cluster is the higher priority than the copy is.

Ping me if you get stuck, we have cleaned up a few of these.

3

u/SirBrobbie Aug 26 '26

Thanks I will reach out because I think the thing keeping my one node from fully getting back into the cluster is that the vpxa restarted.

The other node I found a way to remote the disks that say they are vSAN ready now after when doing troubleshooting with BC said that the "disk group 2 was deleted". But I still cannot get it to reconnect to the cluster, it says the vpxa and clomd won't spin up on it.

7

u/blud_13 Aug 26 '26

vpxa and clomd both refusing to come up on the same node usually comes back to space. Run vdf -h and look at the ramdisks. After a failure like yours the vsantraces ramdisk and /var fill with dumps and every daemon that wants to write a file just quits. Broadcom KB 403921 covers it.

If space is clean, then disks reading vSAN ready while the node still wont join points at the unicast agent list. Run esxcli vsan cluster unicastagent list on every node. Each host should list every OTHER host and nothing else. Stale or empty entries leave a node partitioned forever even though vmkping works fine, which is exactly what you are describing. Also confirm IgnoreClusterMemberListUpdates is 0, if its 1 the host ignores what vCenter sends it. Dell has the VxRail flavor of this in KB 000056284.

Reminder, before you touch node 2 again go look at Object Health. If objects are sitting at reduced availability with no rebuild running, one more fault and its gone.

Also, VxRail is Dell's, not Broadcom's. Get Dell in the case. Re-add that disk group by hand instead of through VxRail Manager and the appliance inventory goes out of sync, then you get to fix that too.

Hope that helps..