Applies to: VMware vSphere 9.x / 8.x / 7.x / 6.7 / 6.5 / 6.0 / 5.5 / 5.1 / 5.0
vSphere High Availability (HA) uses the Fault Domain Manager (FDM) agent on ESXi hosts to monitor host and virtual machine availability and coordinate recovery when a failure occurs. FDM was introduced with vSphere 5 and replaced the older AAM-based HA architecture.
The basic design is still used in current vSphere releases: one host takes the Primary role and the remaining hosts run as Secondary nodes. Network heartbeats and datastore heartbeats help vSphere HA determine whether a host has failed, is isolated, or is part of a network partition.
In This Article
- 1. What Is FDM?
- 2. Primary and Secondary Hosts
- 3. Network Heartbeats
- 4. Datastore Heartbeats
- 5. What Happens If the Primary Host Fails?
- 6. Basic FDM Troubleshooting
- 7. Key Takeaways
1. What Is FDM?
FDM stands for Fault Domain Manager. The FDM agent is deployed on each ESXi host when vSphere HA is enabled for a cluster.
Its job is to monitor cluster members, track protected virtual machines, exchange heartbeat information, and coordinate restart actions when a host or virtual machine failure is detected.
In older vSphere 4.x environments, HA used the AAM agent and a different primary-node model. Starting with vSphere 5.x, HA moved to the FDM architecture.
2. Primary and Secondary Hosts
One ESXi host in the HA cluster is elected as the Primary host. The other hosts operate as Secondary hosts.
- The Primary FDM agent monitors the state of the Secondary hosts and protected virtual machines.
- Secondary FDM agents report host and virtual machine state to the Primary.
- The Primary communicates HA state information to vCenter Server.
This design avoids the older fixed group of primary nodes used by AAM. If the current Primary host becomes unavailable, another host can be elected.
3. Network Heartbeats
FDM agents exchange network heartbeats between the hosts in the HA cluster. If the Primary stops receiving heartbeats from a Secondary host, HA does not immediately assume that the host has failed.
It must first determine whether the host is actually down, network isolated, or separated by a network partition. This distinction matters because the correct HA response depends on the failure type.
Broadcom documents TCP and UDP port 8182 as required for vSphere HA communication between hosts.
4. Datastore Heartbeats
Datastore heartbeating provides another source of information when network communication between HA hosts is lost. vSphere HA uses heartbeat files on shared datastores to help determine whether a host is still alive.
This does not replace normal FDM network communication. Its purpose is to help HA distinguish a real host failure from network isolation or a network partition.
The heartbeat files are stored under the hidden .vSphere-HA directory on the selected datastore.
5. What Happens If the Primary Host Fails?
If the Primary host fails, vSphere HA starts a new election and another eligible host becomes the Primary. The new Primary rebuilds its view of the cluster and continues monitoring the remaining hosts and protected virtual machines.
This election process is a normal part of the FDM architecture. It can also be seen after some vCenter Server upgrades or HA reconfiguration operations when the FDM agent is updated or redeployed.
6. Basic FDM Troubleshooting
When vSphere HA reports an agent, heartbeat, isolation, or partition problem, start with these checks:
- Verify management network connectivity between all ESXi hosts.
- Verify TCP and UDP port 8182 is not blocked between HA hosts.
- Check whether the configured heartbeat datastores are accessible from the affected hosts.
- Review /var/run/log/fdm.log on the ESXi host.
- Use Reconfigure for vSphere HA when an individual host has an HA configuration problem.
7. Key Takeaways
- vSphere HA has used the FDM architecture since vSphere 5.x.
- One host runs as the Primary and the remaining hosts operate as Secondary nodes.
- Network heartbeats are the normal communication path between FDM agents.
- Datastore heartbeats help distinguish host failure from isolation or network partition.
- If the Primary host fails, another eligible host is elected automatically.
References
- Broadcom Knowledge Base – Simulating vSphere High Availability failures
- Broadcom Knowledge Base – vSphere High Availability issues

Cloud and infrastructure professional with nearly two decades of experience in enterprise IT environments, spanning public cloud, private cloud, and hybrid architectures.