Home / Disaster Recovery / Use case: hospital group
Illustrative scenario

Tiered disaster recovery for a hospital group

How we would approach DR for a three-hospital group whose only copy of its data sat in the same building as production.

This is a composite example built from requirements that are common in hospital IT. It is not a specific client engagement, and the figures are design targets for the scenario, not measured results. It shows how we think through a DR problem from start to finish.

Organisation3 hospitals, ~600 beds
Key systemsHIS, lab, PACS imaging
Data~40 TB, mostly images
Pattern usedThree tiers
The situation

One server room, one copy, no tested plan

All clinical systems ran from a server room at the main hospital: the hospital information system (HIS) for admissions, billing and patient records, the lab system, and PACS for X-ray, CT and MRI images. Backups ran every night to a disk appliance in the same room.

A UPS failure took the room down for nine hours. Wards went back to paper, the lab phoned results through, and radiologists could not see previous scans. Nobody could say how long a full restore would take, because a full restore had never been tried.

The management team asked a simple question: if the server room is lost tomorrow, how long until doctors can work normally again, and how much data do we lose?

What we found

Not every system needs the same protection

We sat with the medical director, the lab head, the radiology head and finance, and asked the same two questions about each system: how much data can you afford to lose, and how long can you work without it?

  • HIS and lab were critical. Admissions, medication orders and lab results flow through them all day. Staff could re-enter about 15 minutes of work from paper. Two hours of downtime was the most they could manage safely.
  • PACS was important, but in a specific way. Radiologists needed recent scans quickly. Images older than about three months were rarely needed on the first day.
  • Email, file shares and the remaining servers could wait a day or two.
  • The only backup copy was in the same room as production, so it offered no protection against fire, flood or a long power failure.
The design

Three tiers, one DR site

A colocation rack in another city, far enough away not to share the same power grid, flood risk or local disruption. Each tier uses the cheapest pattern that still meets its targets.

TierSystemsRPO / RTO targetPatternHow
1HIS, lab15 min / 2 hWarm standbyDatabase replication to standby databases at DR. Two application servers run at DR and scale to six on failover.
2PACS1 h / 8 hPilot lightRecent images replicated on a schedule. Older images kept on object storage at DR. Viewer servers started on failover.
3Email, files, others24 h / 48 hBackup & restoreNightly backup copied to immutable storage at DR. Servers restored only when needed.
Tiered DR design: primary hospital data centre replicating to a colocation DR site in another cityDoctors, nurses, lab staffDNS switch on failoverPrimary: main hospital data centreDR: colocation rack, another cityTier 1 · HIS and lab · RPO 15 min · RTO 2 hTier 2 · Imaging (PACS) · RPO 1 h · RTO 8 hTier 3 · Email, files, others · RPO 24 h · RTO 48 hHIS and labapp servers (6)HIS and labdatabasesStandbydatabasesApp servers: 2 on,scale to 6Database replicationasync, lag under 15 minPACS image storageabout 40 TB, growing about 1 TB a monthLast 90 days of images readyolder images on object storageImage replicationevery 15 to 60 minEmail, file shares,other serversBackup server(local copy)Immutable backup copyservers restored when neededNightly backup copy
Solid lines: continuous or scheduled replication. Dashed lines: backup copy, and the user path after failover.
Sizing

How much bandwidth does this need?

Replication bandwidth comes from how much data changes each day, not from total data size. For this scenario:

  • HIS and lab database changes: about 30 GB a day
  • New imaging studies: about 35 GB a day
  • Other backup data sent across: about 10 GB a day

That is about 75 GB a day, which works out to roughly 7 Mbps if spread evenly over 24 hours. Hospital traffic is not spread evenly. If 60% of it arrives during a five-hour morning outpatient peak, the link needs to carry around 20 Mbps at peak without falling behind.

We would recommend two 100 Mbps links from different providers. Each one can carry the peak alone with plenty of headroom, so losing one link does not break the RPO.

The first copy is a separate problem. Sending 40 TB of existing images over a 200 Mbps link would take about 18 days. Shipping encrypted disks to the DR site for the first copy, then replicating only changes over the network, is faster and safer.

Rollout

How it would be delivered

  1. Weeks 1–3: application inventory, dependency mapping, and RPO/RTO sign-off with clinical and finance heads.
  2. Weeks 3–5: DR site selection, network links ordered, bill of materials finalised.
  3. Build: tier 3 backups first (quickest win), then tier 1 replication, then PACS seeding and replication.
  4. Runbooks: one per tier, written with the people who will run them, including the paper downtime procedure for wards.
  5. First drill: isolated failover of tier 1 on a Sunday morning. Measured recovery time and replication lag recorded against the targets, every issue fixed and retested.
What changes

Before and after (design targets)

BeforeAfter (target)How it is checked
HIS data lost if the site is lostUp to 24 hours (since last backup)Up to 15 minutesReplication lag monitored with alerts
HIS back in serviceUnknown. Nine hours in the UPS incidentWithin 2 hoursTimed in each drill
Where the backup copy livesSame room as productionAnother city, plus an immutable copyMonthly restore test
Recovery instructionsIn one engineer's headWritten runbook for each tierTabletop review every quarter
DR running footprintNoneRoughly a third of production computeTier 1 runs at reduced size; tiers 2 and 3 are mostly off
Trade-offs

What we would flag to the hospital

  • Up to 15 minutes of HIS data can still be lost. Replication is asynchronous because the sites are far apart. Wards need a paper downtime form and a clear re-entry step after failover.
  • Older images come back slowly. Scans older than 90 days restore from object storage over hours, not minutes. Radiology agreed this was acceptable for the first day.
  • Database licences matter. Built-in replication features are not included in every database edition. For example, Oracle Data Guard requires Enterprise Edition. If the HIS runs on a lower edition, the replication method and its cost change.
  • DR only stays useful if it is tested. The budget should include two drills a year, or the setup will drift out of date within months.

Not sure where your DR stands?

Send us your application list and a short note on how backups work today. We will come back with a plain gap review: which systems are exposed, what a realistic recovery time looks like, and what it would take to close the gap.