Tiered disaster recovery for a hospital group
How we would approach DR for a three-hospital group whose only copy of its data sat in the same building as production.
This is a composite example built from requirements that are common in hospital IT. It is not a specific client engagement, and the figures are design targets for the scenario, not measured results. It shows how we think through a DR problem from start to finish.
One server room, one copy, no tested plan
All clinical systems ran from a server room at the main hospital: the hospital information system (HIS) for admissions, billing and patient records, the lab system, and PACS for X-ray, CT and MRI images. Backups ran every night to a disk appliance in the same room.
A UPS failure took the room down for nine hours. Wards went back to paper, the lab phoned results through, and radiologists could not see previous scans. Nobody could say how long a full restore would take, because a full restore had never been tried.
The management team asked a simple question: if the server room is lost tomorrow, how long until doctors can work normally again, and how much data do we lose?
Not every system needs the same protection
We sat with the medical director, the lab head, the radiology head and finance, and asked the same two questions about each system: how much data can you afford to lose, and how long can you work without it?
- HIS and lab were critical. Admissions, medication orders and lab results flow through them all day. Staff could re-enter about 15 minutes of work from paper. Two hours of downtime was the most they could manage safely.
- PACS was important, but in a specific way. Radiologists needed recent scans quickly. Images older than about three months were rarely needed on the first day.
- Email, file shares and the remaining servers could wait a day or two.
- The only backup copy was in the same room as production, so it offered no protection against fire, flood or a long power failure.
Three tiers, one DR site
A colocation rack in another city, far enough away not to share the same power grid, flood risk or local disruption. Each tier uses the cheapest pattern that still meets its targets.
| Tier | Systems | RPO / RTO target | Pattern | How |
|---|---|---|---|---|
| 1 | HIS, lab | 15 min / 2 h | Warm standby | Database replication to standby databases at DR. Two application servers run at DR and scale to six on failover. |
| 2 | PACS | 1 h / 8 h | Pilot light | Recent images replicated on a schedule. Older images kept on object storage at DR. Viewer servers started on failover. |
| 3 | Email, files, others | 24 h / 48 h | Backup & restore | Nightly backup copied to immutable storage at DR. Servers restored only when needed. |
How much bandwidth does this need?
Replication bandwidth comes from how much data changes each day, not from total data size. For this scenario:
- HIS and lab database changes: about 30 GB a day
- New imaging studies: about 35 GB a day
- Other backup data sent across: about 10 GB a day
That is about 75 GB a day, which works out to roughly 7 Mbps if spread evenly over 24 hours. Hospital traffic is not spread evenly. If 60% of it arrives during a five-hour morning outpatient peak, the link needs to carry around 20 Mbps at peak without falling behind.
We would recommend two 100 Mbps links from different providers. Each one can carry the peak alone with plenty of headroom, so losing one link does not break the RPO.
The first copy is a separate problem. Sending 40 TB of existing images over a 200 Mbps link would take about 18 days. Shipping encrypted disks to the DR site for the first copy, then replicating only changes over the network, is faster and safer.
How it would be delivered
- Weeks 1–3: application inventory, dependency mapping, and RPO/RTO sign-off with clinical and finance heads.
- Weeks 3–5: DR site selection, network links ordered, bill of materials finalised.
- Build: tier 3 backups first (quickest win), then tier 1 replication, then PACS seeding and replication.
- Runbooks: one per tier, written with the people who will run them, including the paper downtime procedure for wards.
- First drill: isolated failover of tier 1 on a Sunday morning. Measured recovery time and replication lag recorded against the targets, every issue fixed and retested.
Before and after (design targets)
| Before | After (target) | How it is checked | |
|---|---|---|---|
| HIS data lost if the site is lost | Up to 24 hours (since last backup) | Up to 15 minutes | Replication lag monitored with alerts |
| HIS back in service | Unknown. Nine hours in the UPS incident | Within 2 hours | Timed in each drill |
| Where the backup copy lives | Same room as production | Another city, plus an immutable copy | Monthly restore test |
| Recovery instructions | In one engineer's head | Written runbook for each tier | Tabletop review every quarter |
| DR running footprint | None | Roughly a third of production compute | Tier 1 runs at reduced size; tiers 2 and 3 are mostly off |
What we would flag to the hospital
- Up to 15 minutes of HIS data can still be lost. Replication is asynchronous because the sites are far apart. Wards need a paper downtime form and a clear re-entry step after failover.
- Older images come back slowly. Scans older than 90 days restore from object storage over hours, not minutes. Radiology agreed this was acceptable for the first day.
- Database licences matter. Built-in replication features are not included in every database edition. For example, Oracle Data Guard requires Enterprise Edition. If the HIS runs on a lower edition, the replication method and its cost change.
- DR only stays useful if it is tested. The budget should include two drills a year, or the setup will drift out of date within months.
Not sure where your DR stands?
Send us your application list and a short note on how backups work today. We will come back with a plain gap review: which systems are exposed, what a realistic recovery time looks like, and what it would take to close the gap.