Home / Disaster Recovery / HA, fault tolerance and DR
DR guide

High availability, fault tolerance and disaster recovery

The three terms are often used as if they mean the same thing. They do not. Each one protects against a different kind of failure, at a different cost, and a design that has only one of them has a gap. This guide explains the difference in plain terms and shows where each belongs.

The three, side by side

Scroll sideways to see the whole diagram →
HIGH AVAILABILITY: ONE SITEFAULT TOLERANCE: NO INTERRUPTIONDISASTER RECOVERY: SECOND SITELoad balancerhealth checksApp server Arack 1App server Brack 2DatabaseprimaryDatabasereplicaPrimary VMruns the workloadShadow VMrepeats every step, other hostShared storageboth see the same disksPrimary siteyour main data centreDR site, far awaystandby copy, different risk zoneOffline backupspoint in time, immutablesync2lockstep3replication1User or API trafficSynchronous writeData / replication
1. High availability spreads one service across racks or zones in the same site, so a single failed server or rack is absorbed in seconds to minutes. 2. Fault tolerance runs a mirrored copy in lockstep, so a failed host causes no interruption at all, at a high cost per workload. 3. Disaster recovery keeps a copy in a second site far enough away not to share the same fire, flood or power failure, plus backups that survive data corruption.

How they compare

High availability (HA)Fault tolerance (FT)Disaster recovery (DR)
Protects againstA failed server, disk, rack or availability zoneA failed host, with zero interruptionLoss of a whole site: fire, flood, long power failure, ransomware, regional outage
Typical downtimeSeconds to a few minutesNoneMinutes to hours, set by the RTO
Typical data lossNoneNoneZero to a few minutes, set by the RPO
Where the copy livesSame site, different rack or zoneSame site, different hostDifferent site, far enough away to avoid the same event
CostModerate: roughly double the critical componentsHigh: every workload runs twice, often with licence limitsVaries widely with the pattern, from backups only to a full second site
Typical examplesClustered databases, load-balanced app servers, dual power and networkVMware vSphere Fault Tolerance, lockstep hardware, some telecom and control systemsDatabase replication to a second site, cloud DR, backups held off site

Why you usually need all three

Most outages are small: a disk, a server, a switch. High availability handles those without anyone noticing, and that is where it earns its cost. But HA lives inside one site. If the site loses power for a day, every redundant server in it goes down together.

Disaster recovery covers the large, rare event. It is slower and involves a decision to fail over, but it is the only thing that helps when the whole site is gone.

Fault tolerance is the narrow case: a small number of workloads where even a few seconds of interruption is unacceptable. Using it everywhere is expensive and rarely needed.

And none of the three replaces backups. Replication copies corruption and ransomware encryption to the other side within seconds. A backup from yesterday, held where an attacker cannot delete it, is what lets you go back to before the problem started.

Four common mistakes

Calling HA “our DR”

Two clustered servers in the same room are high availability. They share the same building, power and network, so one event takes both down.

Treating replication as backup

Replication copies mistakes as faithfully as data. A deleted table or encrypted volume reaches the DR site almost at once.

Fault tolerance for everything

Running every workload twice in lockstep doubles cost for systems that could tolerate a minute of interruption. Keep it for the few that cannot.

A DR site too close

A DR site across the road shares the same flood plain, substation and fibre routes. Distance should match the risks you are protecting against.

Not sure where your DR stands?

Send us your application list and a short note on how backups work today. We will come back with a plain gap review: which systems are exposed, what a realistic recovery time looks like, and what it would take to close the gap.