Home / Whitepapers / Data centre and disaster recovery
Whitepaper · Data centre and disaster recovery

Disaster Recovery That Works on the Day

Choosing, building and proving a DR strategy for enterprise and government IT. Written for CIOs, IT heads, infrastructure architects, and risk and compliance teams.

Reading time22 minutes
LengthApprox. 16 pages
Editionv1.0, October 2026
FormatWeb and PDF

What is inside

  • How to set RPO and RTO per application tier, from business impact rather than guesswork
  • The five DR patterns compared on recovery time, data loss and running cost
  • A four-question decision framework, and why most organisations need three tiers
  • Bandwidth, storage and cost formulas with a worked example
  • Why replication is not ransomware protection, and what is
  • A testing programme that turns DR from a document into evidence
  • A 40-point readiness checklist you can use tomorrow
Or start reading online ↓
Download the full PDFApprox. 16 pages, including the 40-point checklist

We confirm your email with a one-time code, then the PDF downloads straight away. No newsletters unless you ask.

01

Executive summary

Most organisations have a disaster recovery plan. Far fewer have disaster recovery that works when it is needed. The gap is rarely the technology. It is that recovery targets were never agreed with the business, every application was protected the same way, backups shared the same failure as production, and the plan was never rehearsed end to end.

This paper sets out a practical way to close that gap. It is written for the people who have to make DR decisions and defend them: CIOs and IT heads, infrastructure architects, and the risk and compliance teams who sign off.

Five things to take away

  1. Set targets per tier, not per company. Agree recovery time (RTO) and acceptable data loss (RPO) for each group of applications, starting from what an hour of downtime actually costs.
  2. Use two or three DR patterns, not one. A handful of critical systems need a warm or hot copy. Most applications do not. Tiering typically cuts DR running cost by 40 to 60 percent against protecting everything the same way.
  3. Replication is not backup. Replication copies ransomware and mistakes as fast as good data. You need immutable, versioned backups that attackers cannot reach.
  4. The runbook is part of the architecture. Most of the time lost in a real disaster is spent deciding, finding people and fixing what was never written down.
  5. Only a tested recovery counts. Measure RPO and RTO in drills, report them against targets, and fix what breaks. Untested DR is a guess.
02

Why DR fails in practice

Power failures, storage faults, network outages, fire, flooding, ransomware and plain human error all take data centres down. None of these are rare across a five-year horizon. The cost shows up as lost revenue, penalties, customer dissatisfaction, compliance findings and damage to reputation, and it grows with every hour of downtime.

When we review DR set-ups, the same failure modes appear again and again:

What we findWhat happens on the day
Targets nobody agreedIT built for 4 hours; the business assumed 30 minutes. The argument happens during the outage.
Everything protected the same wayEither the budget runs out before critical systems are covered, or money is wasted keeping test servers hot.
Backups in the same failure domainBackup server in the same room, on the same power feed, or reachable with the same admin password as production.
Replication treated as backupRansomware-encrypted files replicate to DR within minutes. Both sites are now encrypted.
Only the data is copiedVolumes are at DR, but firewall rules, certificates, DNS, licence servers and identity are not.
No one decidesThere is no named person or agreed criteria for declaring a disaster. Hours pass before failover starts.
Never tested end to endThe first full failover happens during a real disaster, and it is also the first time failback is attempted.
In short: DR fails at the seams between business, technology and people. A good design addresses all three.
03

RPO and RTO, set properly

Recovery Point Objective (RPO) is the maximum amount of data, measured in time, the business can afford to lose. If data is replicated every five minutes, up to five minutes of work can be lost. Recovery Time Objective (RTO) is how long a service can be unavailable before the impact becomes unacceptable. An RPO of 15 minutes and an RTO of 2 hours means: lose at most the last 15 minutes of work, and be running again within 2 hours.

Two rules of thumb follow. A lower RPO generally requires more replication and more bandwidth. A lower RTO generally requires more infrastructure running at the DR site. Both cost money, so the targets must come from business impact, not from preference.

A short business impact analysis

For each application, ask its business owner four questions and write the answers down:

  • What does one hour of downtime cost? Revenue, penalties, patient or citizen impact, regulatory exposure.
  • How much recent data can be re-entered by hand from paper, email or another system?
  • Does a regulator, contract or customer SLA set a number?
  • What else depends on this system, and what does it depend on?

Then group applications into tiers. Three tiers are enough for most organisations:

TierTypical systemsRPO targetRTO target
Tier 1: criticalCore banking, payments, ERP transactions, hospital information system, citizen-facing portalsNear zero to 15 minutes15 minutes to 2 hours
Tier 2: importantCustomer portals, MIS, email, document management, HR and payroll15 minutes to 4 hours4 to 12 hours
Tier 3: deferrableDevelopment and test, archives, analytics sandboxes, internal tools24 hours1 to 3 days
Watch the dependencies. A tier 1 application that relies on a tier 3 directory or licence server is really a tier 3 application. Map dependencies before you assign tiers.
04

The five DR patterns

There are five common ways to set up disaster recovery. They differ in what is already running at the DR site before a disaster, and therefore in recovery time, data loss and cost.

PatternWhat is ready at DRTypical RPOTypical RTODR running cost vs production
Backup and restoreBackup copies only. Servers are rebuilt after the disaster.Hours (time since last backup)Hours to daysAbout 10 to 20%
Pilot lightData replicated. Core services such as database and directory run small. Application servers are off.MinutesA few hoursAbout 20 to 35%
Warm standbyThe full stack runs at reduced size and is scaled up on failover.Seconds to minutesUnder an hour to a few hoursAbout 40 to 60%
Active-passiveA full-size copy is ready to take over.Near zero (synchronous) to minutesMinutesAbout 80 to 100%
Active-activeBoth sites serve users at the same time.Near zeroNear zero for users100% or more, plus application changes

Backup and restore

Scheduled backups or volume snapshots are copied to storage at another site, ideally object storage with immutability. After a disaster, infrastructure is rebuilt and data restored. Cheap to run, slow to recover, and the RPO is the time since the last good backup. Right for tier 3 and for long-term retention behind every other pattern.

Pilot light

The data layer is kept live at the DR site: database replication, for example asynchronous log shipping or a Galera or Always On replica, and storage mirroring on a schedule. Application and web servers exist only as templates and are started on failover. Recovery takes hours because servers must boot, attach data and be checked, but running cost stays low.

Warm standby

A scaled-down copy of the whole stack runs continuously at DR, with continuous block or database replication. On failover it is scaled up and traffic is redirected. Because everything is already running, it can be tested without disruption, which is the pattern’s biggest practical advantage.

Active-passive

A full-size copy stands ready. With synchronous replication, every write is confirmed at both sites, which gives near-zero data loss but limits the distance between sites to roughly 100 km and adds write latency. Network gateways and load balancers are pre-configured so failover is a decision, not a build.

Active-active

Both sites serve users at the same time behind a global load balancer, and the loss of one is close to invisible. It is the most expensive pattern and usually needs application changes, because two sites writing to the same data must avoid conflicts and split-brain. Worth it only for services designed for it from day one.

Taking this to a meeting? Get the PDF to share with your team: same content, printable checklist, no clutter.
Download the PDF
05

Choosing a strategy: a four-question framework

The choice of pattern follows from four questions, asked per tier:

  1. What does an hour of downtime cost? If nobody can put a number or a consequence on it, the system probably does not need the expensive patterns.
  2. How much recent data can be re-entered? If staff can re-key the last 30 minutes, asynchronous replication is usually enough and synchronous replication is not needed.
  3. Does a regulator or contract set the number? Banks, insurers, market intermediaries and many government systems have written continuity requirements. Start from those.
  4. Can the application be changed? If not, active-active is off the table, and active-passive or warm standby is the realistic ceiling.

Most organisations land on the same shape: active-passive or warm standby for a handful of critical systems, pilot light for important business applications, and backup and restore for everything else, with immutable backups behind all three.

SectorTypical tier 1 patternWhat drives it
Banking and financial servicesActive-passive, sometimes three sitesRound-the-clock transactions, regulator DR drills, near-zero data loss
Government servicesWarm standby or active-passiveCitizen services during emergencies, data residency, audit trail
HealthcareWarm standbyClinical systems need fast recovery; records must stay on-premises
ManufacturingPilot light, with warm standby for MESPlant continuity at balanced cost; plants often keep running locally
AI and GPU platformsBackup of data and checkpoints; warm standby for inferenceGPUs are too costly to duplicate idle; training restarts from checkpoints
06

Reference architecture

The diagram shows the tiered design most organisations end up with. Users reach applications through a global DNS or load balancer with short time-to-live values, so traffic can be moved in minutes. Tier 1 runs warm at the DR site, tier 2 is a pilot light, and tier 3 relies on backups.

USERS AND TRAFFIC STEERINGPRIMARY DATA CENTREDR SITEBACKUP VAULTUsers and partnersbranches, staff, customersGlobal DNS or load balancershort TTL, health checksTier 1: critical appscore banking, ERP, HISTier 2: business appsportals, MIS, emailTier 3: other systemsdev, archives, toolsDatabasesclustered, logs shippedPrimary storagevolumes, files and objects, with snapshotsTier 1 warmhalf size, runningTier 2 pilottemplates, offDatabase replicaasynchronous, minutes behindStorage replicasnapshot mirroring on a scheduleRunbookstested every quarterClean restoreisolated and scannedImmutable backupsobject lock, 35 daysHTTPSon failoverrestore test123User or API trafficControl / API callData / replicationScheduled copy

Numbered flows: (1) database changes replicate asynchronously to the DR replica, (2) storage is mirrored on a schedule, with a shorter interval for tier 1 volumes, (3) immutable backups are written nightly to a vault that production credentials cannot reach. Restores from the vault are tested regularly in an isolated area.

What must exist at the DR site besides data

  • Identity: a directory controller and the same roles and groups, so people can log in.
  • Network: firewall rules, VPN and partner links, IP plan or DNS changes, all kept as code.
  • Certificates and secrets: renewed at both sites on the same schedule.
  • Licences: licence servers and any standby licence entitlements, checked against contracts.
  • Monitoring and logging: so the team can see what is happening after failover.
  • Runbooks and contact lists: stored somewhere reachable when the primary site is not.
On OpenStack and Ceph: site-level DR is built from Ceph RBD mirroring between clusters (snapshot-based for most volumes, journal-based for the few that need a lower RPO), images published to both sites from one pipeline, projects and networks recreated from Heat or Terraform, and Ansible playbooks that promote volumes and start instances in dependency order. High availability inside one site, such as live migration or instance HA, is not DR.
07

Sizing and cost

Replication bandwidth

Bandwidth is driven by how much data changes, not by how much data exists. A useful first estimate:

Required bandwidth (Mbps) = daily change (GB) × 8,000 ÷ seconds in the busy window × 1.3 headroom

For example, 400 GB of changes a day, most of it in a 10-hour business window, needs about 400 × 8,000 ÷ 36,000 × 1.3, or roughly 115 Mbps of sustained replication bandwidth. Measure change rates from storage or database logs over at least two weeks, including month end, before buying links. Provision two links from different providers.

Storage

  • DR replica: at least one full copy of protected data, plus room for snapshots at the DR site.
  • Backups: retention multiplied by daily change, after deduplication. A 30-day retention commonly needs 1.5 to 3 times the protected data.
  • Seeding: tens of terabytes can take weeks over a WAN link. Shipping encrypted disks for the first copy is often faster.

A worked example

Production has 40 virtual machines and 30 TB of data. After business impact analysis, 8 VMs are tier 1, 20 are tier 2 and 12 are tier 3.

Option A: everything active-passiveOption B: protected by tier
What runs at DR every dayAll 40 VMs at full sizeAbout 7 VM-equivalents: tier 1 at half size, tier 2 databases small
Tier 1 recoveryMinutesSame: minutes
Tier 2 recoveryMinutesA few hours
Tier 3 recoveryMinutesOne to two days from backup
DR running cost vs productionAbout 80 to 100%About 35 to 50%

Where the DR site is on-premises or in colocation, hardware for tiers 1 and 2 must still be ready, because servers cannot be bought on the day; the saving is in power, licences and tier 3 capacity. In cloud DR, full capacity is created only on drill or disaster day, which is where tiering saves the most.

Costs that are often missed

  • Failback: moving data back after a cloud failover can carry large egress charges.
  • Licences: some vendors charge in full for a standby database or hypervisor. Read the contract.
  • Drills: a full failover test takes several people most of a day, every year.
  • Keeping DR current: patching templates, renewing certificates and updating firewall rules at DR is ongoing work.
08

Ransomware changes the design

Classic DR assumes the primary site is lost but the data is good. Ransomware breaks that assumption: the data itself is the problem, and attackers deliberately go after backups and DR systems first. Replication makes it worse, because encrypted files replicate to the DR site within minutes.

A design that survives ransomware adds four things:

ControlWhy it matters
Immutable backupsObject lock or write-once storage, so backups cannot be changed or deleted for the retention period, even by an administrator.
Separate credentials and managementBackup and DR systems use their own accounts and multi-factor authentication, not the production directory.
An isolated copyAt least one copy in a vault that opens only for scheduled replication windows, or offline media.
Clean restore points and a clean roomBackups are scanned, and restores happen first into an isolated environment where identity and critical apps are rebuilt and checked before reconnecting.
Plan the order: after ransomware, identity usually has to be rebuilt first, then core infrastructure services, then tier 1 applications. Write that order into the runbook and rehearse it.
09

Testing: turning a plan into evidence

DR that has never been tested is a guess. Most problems appear the first time someone actually fails over: a missing firewall rule, an expired certificate, a licence server that only exists in production. Testing finds these while there is no real disaster.

TestWhat happensSuggested frequency
Tabletop walk-throughThe team reads the runbook step by step and checks every step, name and phone number.Every quarter
Restore testRestore one server or database from backup and confirm it works.Every month, rotating systems
Isolated failoverBring a tier up at DR on an isolated network, without touching production users.Twice a year for each tier
Full failover and failbackMove real users to DR for an agreed window, then move back.Once a year for tier 1, or as a regulator requires

On drill day

  • Before: freeze changes, confirm replication lag is within RPO, tell users and the helpdesk, and write down what success means.
  • During: time every step, record every manual fix however small, have the application owner test, and agree a hard stop for rolling back.
  • After: report measured RPO and RTO against targets, turn every issue into a fix with an owner and a date, update the runbook that week, and book the next drill.

Keeping DR working between drills

  • Alert on replication lag. Lag that quietly grows past the RPO is the most common silent failure.
  • Add “and at DR?” to every change: new applications, firewall changes, certificate renewals.
  • Patch DR on the same schedule as production, and upgrade both sites together.
  • Review tiers once a year. Last year’s tier 3 system may now be critical.
10

A practical roadmap

For a mid-sized environment, a DR programme typically runs in four stages. Procurement of links, hardware and DR site space usually sets the pace.

StageTypical durationOutcome
1. Assess3 to 6 weeksApplication inventory, dependencies, business impact analysis, agreed tiers and targets, gap report.
2. Design3 to 4 weeksPattern per tier, reference architecture, sizing, bandwidth, cost model, runbook outline.
3. Build6 to 16 weeksDR site, links, replication, backup vault, automation, monitoring of replication lag.
4. Prove2 to 4 weeks, then ongoingFirst isolated failover per tier, first full drill for tier 1, evidence pack, test calendar.
11

Ten common mistakes

  1. One RPO and RTO for the whole company. Either too expensive or not protective enough.
  2. Choosing a product before agreeing targets. The product then defines the targets.
  3. Treating replication as backup. It copies ransomware and mistakes just as fast.
  4. DR site sharing risks with production. Same building, power feed, flood zone or admin accounts.
  5. Copying data but not configuration. Networks, certificates and licences rebuilt by hand under pressure.
  6. No named decision-maker. Hours lost before anyone declares the disaster.
  7. Long DNS time-to-live values. Failover is done but users keep going to the dead site for hours.
  8. Never testing failback. Getting back is often the harder half.
  9. Not monitoring replication lag. The RPO is silently missed for weeks.
  10. Letting the two sites drift. Different software versions cause surprises on failover day.
12

DR readiness checklist (40 points)

Use this list to score your current position. Anything you cannot tick with evidence is a gap worth closing.

Governance

  • A named owner for DR and a named person who can declare a disaster
  • Agreed criteria for declaring a disaster
  • RPO and RTO agreed and signed off per tier by business owners
  • Regulatory and contractual DR requirements listed
  • DR included in change management (“and at DR?”)
Applications and data 5 points Infrastructure 8 points Backup and ransomware 6 points Runbooks and people 5 points Testing and evidence 6 points Operations 5 points

The remaining 35 points are in the PDF, laid out as a printable checklist.

Get the full checklist
13

Glossary

TermMeaning
RPORecovery Point Objective: maximum acceptable data loss, measured in time.
RTORecovery Time Objective: maximum acceptable downtime for a service.
Synchronous replicationEach write is confirmed at both sites before it completes. Near-zero data loss, distance-limited.
Asynchronous replicationWrites are copied to the second site shortly after. Small possible data loss, any distance.
Failover / failbackMoving service to the DR site, and later moving it back.
Immutable backupA backup that cannot be changed or deleted for its retention period.
Split-brainBoth sites believe they are primary and accept writes, causing conflicting data.
Tabletop exerciseA walk-through of the runbook without touching systems.

About Vakratron Systems

Vakratron Systems is a vendor-neutral infrastructure design firm. We design data centre, disaster recovery, cloud, GPU and AI platforms for enterprises and government buyers, write our assumptions down, and stay with a design until it is running and tested.