A recovery vault attackers cannot reach, and a clean room to rebuild from it
A listed pharmaceutical company with three plants and a research centre asked a hard question after a peer was hit by ransomware: if every server and every backup we can reach from the network were encrypted tomorrow, what would we rebuild from, and where? This design answers it with an isolated vault that opens for two hours a night, copies that are scanned before they are trusted, and a clean room where the company can rebuild its identity system and its regulated applications.
Good backups, all reachable by the same attacker
The company runs ERP, a laboratory information system, quality management, manufacturing execution and an electronic document system across three plants and a research centre. Many of these are validated computerised systems under GMP rules, and batch release depends on them. Backups are taken nightly to an appliance with immutable snapshots and copied to a second site.
A tabletop exercise exposed the gap. The backup appliance, the second site and the virtualisation platform all trust the same Active Directory. An attacker who takes control of AD, as most ransomware crews now do before encrypting anything, could reach the backup console, shorten retention or wait until the immutable window expires. Nobody could say which restore point was clean, or where to rebuild AD itself.
The board’s risk committee asked for a recovery capability that does not depend on anything an attacker could reach from the production network, with a tested recovery time, and with evidence that restored GxP systems can be trusted before batch release restarts.
What could not be compromised
- The vault must share no credentials, no directory and no management plane with production.
- Recovery of tier-0 (AD, DNS, PKI) within 24 hours and of batch-release systems within 72 hours.
- Restored GxP systems must show data integrity and validated state before any regulated decision is made on them.
- Incident handling must meet CERT-In’s 6-hour reporting and the company’s SEBI disclosure obligations as a listed entity.
- No change to production applications, and no extra load on plants during the working day.
Four ways to protect the last copy, compared
Each option was tested against one scenario: an attacker with domain administrator rights for three weeks before encrypting.
The vault pulls copies in, checks them, and stays closed the rest of the day
The vault sits in its own locked room with its own network, its own administrators and its own key and time sources. The link to production is physically off for 22 hours a day. Vault control, from inside, opens it for a two-hour window, pulls the night’s copies, closes it, then scans what arrived. Next to the vault, a clean room holds enough compute to rebuild the company’s minimum working core.
| Building block | Why it is there |
|---|---|
| 1 Timed air gap | A single replication link whose switch ports are physically disabled outside the window. The vault opens it from inside; nothing in production can open it, and no production account exists inside the vault. |
| 2 Vault copies with retention lock | Deduplicated copies of the critical systems, kept 90 days under a compliance-mode retention lock that cannot be shortened by any administrator. The vault uses its own time source so the lock cannot be tricked by a changed clock. |
| 3 Backup scanning | Every new copy is checked for known malware, sudden changes in file entropy that point to encryption, mass deletions and changes to AD that should not be there. The last copy that passes is recorded as the clean restore point for each system. |
| 4 Vault control | Runs the window, the scans and the health checks, all from inside the vault. Its logs leave only on a one-way export to the SOC, so the SOC sees vault health without having any way in. |
| 5 Clean room | An isolated set of recovery hosts and flash storage, sized for the minimum working core of the company: identity, ERP, LIMS, QMS, eDMS and batch release. It never connects to production until systems are declared clean. |
| 6 Gold images and clean-room AD | Hardened, versioned builds of operating systems, domain controllers and key applications. AD is rebuilt from the last clean copy into the clean room first, credentials are reset, and only then are applications brought up. |
| 7 Recovery runbooks | Step-by-step, ordered by dependency, with named owners from IT, quality assurance and application teams. Most steps are automated and timed in every test. |
| 8 Validation checks | Pre-agreed checks for each GxP system: audit trail continuity, record counts, checksums against the last known good state, and quality assurance sign-off before batch release resumes. |
Sized around the minimum working company, not every server
Business impact workshops identified about 120 virtual machines across 14 systems that the company needs to release a batch, pay suppliers and staff, and meet regulators. Everything else recovers later from normal backups once the core is clean.
| Item | Figure | Basis |
|---|---|---|
| Protected scope | ~180 TB across 14 systems | Tier 0 and tier 1 only, from business impact analysis |
| Nightly change sent to vault | ~2.4 TB | ~4% daily change = 7.2 TB, deduplicated about 3:1 |
| Replication window | 2 hours a night on 10G | 2.4 TB at ~1 GB/s effective is ~40 minutes; 3x headroom for month end |
| Vault capacity | ~500 TB usable | 180 TB base + 90 days x 2.4 TB = ~400 TB, plus 25% growth |
| Scanning throughput | ~2.4 TB a night, full rescan weekly | Changed data scanned nightly; full copy of each system scanned every 7 days |
| Clean-room compute | 8 hosts, ~1,000 vCPU, 6 TB memory | 120 VMs at measured production size, no overcommit on memory |
| Clean-room storage | ~200 TB flash | Restore at ~2 GB/s puts 180 TB back in about a day, tier 0 in hours |
Vault, scanning, clean room and network are estimated at ₹12 to 16 crore, with annual support and testing effort on top. Recovery times in the outcomes are confirmed by the first full recovery test, not assumed.
Five stages, ending in a real recovery test
The vault is itself a system that GxP data passes through, so it is designed, qualified and tested with quality assurance from the start, not after.
Assess and scope
Weeks 1 to 6
Business impact analysis, dependency maps for the 14 core systems, recovery order and targets agreed with the business and quality assurance.
Gate: Risk committee approves scope, recovery targets and budget.
Build vault and clean room
Weeks 5 to 16
Room, network, vault storage, scanning, vault control and clean-room hosts built with separate identity and keys. Gold images created.
Gate: Independent penetration test finds no path from production into the vault.
Qualify
Weeks 14 to 20
Installation and operational qualification of the vault and clean room under the company’s computerised system validation procedure.
Gate: Quality assurance signs the qualification report.
First full recovery test
Weeks 20 to 24
AD, ERP, LIMS and QMS rebuilt in the clean room from vault copies, then checked by application owners and QA. Every step timed.
Gate: Tier 0 in under 24 hours, batch-release systems in under 72 hours.
Operate and rehearse
Ongoing
Daily windows and scans, monthly restore of one system, quarterly full test of a recovery scenario, yearly test with the board’s crisis team.
Gate: Each test report reviewed by the risk committee.
What could defeat a vault, and how the design closes each gap
| Risk | What could happen | How the design handles it |
|---|---|---|
| Dormant malware in backups | Attackers sit quietly for weeks so infected data is copied into the vault | Scanning on every copy, 90-day retention so restore points older than the intrusion exist, and AD checks for changes that should not be there. |
| Insider or stolen vault credentials | Someone shortens retention or deletes copies | Compliance-mode lock that no one can override, separate admin accounts with hardware tokens, two-person rule for any vault change. |
| Window left open | The link stays up and becomes a path in | Ports disabled by default; an alarm and automatic shutdown if the link is up outside the window. |
| Recovery takes longer than planned | Untested dependencies slow the rebuild | Dependency-ordered runbooks, timed every quarter, with the slowest steps automated first. |
| Restored data not trusted by QA | Batch release cannot restart even when systems are back | Validation checks agreed in advance, rehearsed with QA in every test, and audit trail continuity checked automatically. |
What was optimised
Closed by default
The link is physically off for most of the day, and only the vault can open it.
Systems in scope
Only the minimum working company goes into the vault, which keeps cost and recovery time down.
Deduplicated copies
Only changed, unique data crosses each night, so a 10G link and a short window are enough.
Scanned before trusted
A clean restore point is a measured fact, not a guess made during a crisis.
Shared credentials
Separate directory, keys and time source, so compromising production gives no way in.
Tested recovery
Recovery times come from real rebuilds, with QA involved every time.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Copy an attacker cannot reach | None | Vault copies with no shared admin path, locked for 90 days |
| Tier-0 recovery (AD, DNS, PKI) | Unknown, never tested | Under 24 hours |
| Batch-release systems back | Weeks, by estimate | 48 to 72 hours, validated |
| Knowing which copy is clean | Guesswork during an incident | Recorded for each system every night |
| Data loss for core systems | Up to 24 hours, if backups survived | Up to 24 hours, guaranteed copy available |
| Evidence for board and regulators | Backup job reports | Quarterly recovery test reports with timings and QA sign-off |
Targets are set in the scoping stage and confirmed by the first full recovery test. Recovery times depend on the size of the core scope and on how much of the runbook is automated.
What a team needs to deliver this
Cyber recovery architecture
Vault isolation, timed air gaps, retention locks and separate identity and key domains.
Backup and storage design
Deduplication, replication windows, capacity planning and restore throughput.
Active Directory recovery
Forest recovery, tier-0 hardening and credential resets in an isolated environment.
Malware analysis of backups
Scanning, entropy analysis and choosing clean restore points with evidence.
GxP computerised system validation
Qualification of recovery infrastructure and data integrity checks QA can sign.
Recovery testing and runbooks
Dependency-ordered runbooks, automation and quarterly tests that produce board-ready evidence.
Other scenarios
Three-site core banking DR
A small finance bank protects its core banking and payment switch across three sites, with zero data loss to a near DR site and a far site in another seismic zone.
Read →Disaster recoveryMulti-tenant DR as a Service
A regional data-centre operator builds a shared DR service for 40 to 60 mid-size customers, with isolated tenant networks, self-service testing and billing per protected VM.
Read →Security operations and DRSOC infrastructure across DC and DR
A private bank builds its own Security Operations Centre across two data centres, sized from 60,000 events per second, with a DR SIEM that can take over.
Read →Could you rebuild if every backup you can reach today were encrypted?
Tell us which systems keep your business running and how your backups are managed today. We will come back with a plain view of where an attacker could reach your last copy, what a vault and clean room would need, and how quickly you could realistically recover.