High availability, fault tolerance and disaster recovery
The three terms are often used as if they mean the same thing. They do not. Each one protects against a different kind of failure, at a different cost, and a design that has only one of them has a gap. This guide explains the difference in plain terms and shows where each belongs.
The three, side by side
How they compare
| High availability (HA) | Fault tolerance (FT) | Disaster recovery (DR) | |
|---|---|---|---|
| Protects against | A failed server, disk, rack or availability zone | A failed host, with zero interruption | Loss of a whole site: fire, flood, long power failure, ransomware, regional outage |
| Typical downtime | Seconds to a few minutes | None | Minutes to hours, set by the RTO |
| Typical data loss | None | None | Zero to a few minutes, set by the RPO |
| Where the copy lives | Same site, different rack or zone | Same site, different host | Different site, far enough away to avoid the same event |
| Cost | Moderate: roughly double the critical components | High: every workload runs twice, often with licence limits | Varies widely with the pattern, from backups only to a full second site |
| Typical examples | Clustered databases, load-balanced app servers, dual power and network | VMware vSphere Fault Tolerance, lockstep hardware, some telecom and control systems | Database replication to a second site, cloud DR, backups held off site |
Why you usually need all three
Most outages are small: a disk, a server, a switch. High availability handles those without anyone noticing, and that is where it earns its cost. But HA lives inside one site. If the site loses power for a day, every redundant server in it goes down together.
Disaster recovery covers the large, rare event. It is slower and involves a decision to fail over, but it is the only thing that helps when the whole site is gone.
Fault tolerance is the narrow case: a small number of workloads where even a few seconds of interruption is unacceptable. Using it everywhere is expensive and rarely needed.
And none of the three replaces backups. Replication copies corruption and ransomware encryption to the other side within seconds. A backup from yesterday, held where an attacker cannot delete it, is what lets you go back to before the problem started.
Four common mistakes
Calling HA “our DR”
Two clustered servers in the same room are high availability. They share the same building, power and network, so one event takes both down.
Treating replication as backup
Replication copies mistakes as faithfully as data. A deleted table or encrypted volume reaches the DR site almost at once.
Fault tolerance for everything
Running every workload twice in lockstep doubles cost for systems that could tolerate a minute of interruption. Keep it for the few that cannot.
A DR site too close
A DR site across the road shares the same flood plain, substation and fibre routes. Distance should match the risks you are protecting against.
Related guides
DR layer by layer
Traffic, network, databases, storage and messaging, with on-premises and cloud equivalents.
Read the guide →Choosing a DR pattern
Backup and restore, pilot light, warm standby, active-passive and active-active.
Read the guide →Testing DR
Drill types, how often to run them, and a drill-day checklist.
Read the guide →Not sure where your DR stands?
Send us your application list and a short note on how backups work today. We will come back with a plain gap review: which systems are exposed, what a realistic recovery time looks like, and what it would take to close the gap.