Home / Deployment scenarios / Cyber recovery vault for ransomware
Reference deployment Cyber recovery

A recovery vault attackers cannot reach, and a clean room to rebuild from it

A listed pharmaceutical company with three plants and a research centre asked a hard question after a peer was hit by ransomware: if every server and every backup we can reach from the network were encrypted tomorrow, what would we rebuild from, and where? This design answers it with an isolated vault that opens for two hours a night, copies that are scanned before they are trusted, and a clean room where the company can rebuild its identity system and its regulated applications.

SectorPharmaceuticals, listed
Protected scope~180 TB critical data
Programme6 months to first full test
ModelOn-premises vault and clean room
The situation

Good backups, all reachable by the same attacker

The company runs ERP, a laboratory information system, quality management, manufacturing execution and an electronic document system across three plants and a research centre. Many of these are validated computerised systems under GMP rules, and batch release depends on them. Backups are taken nightly to an appliance with immutable snapshots and copied to a second site.

A tabletop exercise exposed the gap. The backup appliance, the second site and the virtualisation platform all trust the same Active Directory. An attacker who takes control of AD, as most ransomware crews now do before encrypting anything, could reach the backup console, shorten retention or wait until the immutable window expires. Nobody could say which restore point was clean, or where to rebuild AD itself.

The board’s risk committee asked for a recovery capability that does not depend on anything an attacker could reach from the production network, with a tested recovery time, and with evidence that restored GxP systems can be trusted before batch release restarts.

What could not be compromised

  • The vault must share no credentials, no directory and no management plane with production.
  • Recovery of tier-0 (AD, DNS, PKI) within 24 hours and of batch-release systems within 72 hours.
  • Restored GxP systems must show data integrity and validated state before any regulated decision is made on them.
  • Incident handling must meet CERT-In’s 6-hour reporting and the company’s SEBI disclosure obligations as a listed entity.
  • No change to production applications, and no extra load on plants during the working day.
Options weighed

Four ways to protect the last copy, compared

Each option was tested against one scenario: an attacker with domain administrator rights for three weeks before encrypting.

OptionWhat worksWhat does notVerdict
Rely on immutable snapshots in the current backup systemAlready in place, no new spend.Same directory and console as production. Retention can be attacked through support channels or by waiting.Rejected as sole answer
Cloud backup copy with object lockOff-site and immutable.Restoring ~180 TB over the WAN takes days, and the clean rebuild still needs somewhere to happen.Rejected
Offline tape rotated off-siteTruly offline and cheap per TB.Slow to restore, hard to scan for malware, and recovery takes weeks for critical systems.Kept for long-term archive
Isolated vault with timed air gap, scanning and clean roomNo shared admin path, scanned restore points, a place to rebuild, and a recovery time that can be tested.Higher cost, and recovery tests need business time every quarter.Chosen
Target architecture

The vault pulls copies in, checks them, and stays closed the rest of the day

The vault sits in its own locked room with its own network, its own administrators and its own key and time sources. The link to production is physically off for 22 hours a day. Vault control, from inside, opens it for a two-hour window, pulls the night’s copies, closes it, then scans what arrived. Next to the vault, a clean room holds enough compute to rebuild the company’s minimum working core.

Scroll sideways to see the whole diagram →
PRODUCTION DATA CENTRETIMED AIR GAPCYBER RECOVERY VAULTown admins, own keys, own clockCLEAN ROOMisolated rebuild and checksCritical applicationsERP, LIMS, QMS, MES, eDMSPrimary backupimmutable copies, 35 daysActive Directorysingle forest, ~4,500 usersEDR and NDRendpoints, servers, networkSOC and SIEM24 x 7, incident leadTimed linkopen 2 h a nightNo admin pathseparate loginsVault controlopens link, runs jobs, logsClean restore pointslast verified copy per systemVault copiesretention lock, 90 daysBackup scanningmalware, entropy, file changesRecovery hostsisolated compute and storageGold imagesOS, AD and app buildsRestored applicationsERP, LIMS, QMS firstClean-room ADforest rebuilt, tier 0 firstRecovery teamIT, QA and app ownersValidation checksGxP and data integrity1nightlysystem stateopens windowverified cleanalerts4invoke recoveryrestorerebuildrunbooks6GxP checks235Data / replicationScheduled copyControl / API callException / alert
Numbered flows: (1) production systems are backed up nightly to immutable primary backup, (2) during the window the vault pulls a copy across the timed link, which vault control then closes, (3) every copy is scanned and the last clean one per system is recorded, (4) in an incident the SOC invokes the recovery team, (5) clean copies restore into isolated recovery hosts, where AD is rebuilt from gold images first, (6) restored GxP applications pass validation checks before release to the business.
Building blockWhy it is there
1 Timed air gapA single replication link whose switch ports are physically disabled outside the window. The vault opens it from inside; nothing in production can open it, and no production account exists inside the vault.
2 Vault copies with retention lockDeduplicated copies of the critical systems, kept 90 days under a compliance-mode retention lock that cannot be shortened by any administrator. The vault uses its own time source so the lock cannot be tricked by a changed clock.
3 Backup scanningEvery new copy is checked for known malware, sudden changes in file entropy that point to encryption, mass deletions and changes to AD that should not be there. The last copy that passes is recorded as the clean restore point for each system.
4 Vault controlRuns the window, the scans and the health checks, all from inside the vault. Its logs leave only on a one-way export to the SOC, so the SOC sees vault health without having any way in.
5 Clean roomAn isolated set of recovery hosts and flash storage, sized for the minimum working core of the company: identity, ERP, LIMS, QMS, eDMS and batch release. It never connects to production until systems are declared clean.
6 Gold images and clean-room ADHardened, versioned builds of operating systems, domain controllers and key applications. AD is rebuilt from the last clean copy into the clean room first, credentials are reset, and only then are applications brought up.
7 Recovery runbooksStep-by-step, ordered by dependency, with named owners from IT, quality assurance and application teams. Most steps are automated and timed in every test.
8 Validation checksPre-agreed checks for each GxP system: audit trail continuity, record counts, checksums against the last known good state, and quality assurance sign-off before batch release resumes.
Sizing, worked out

Sized around the minimum working company, not every server

Business impact workshops identified about 120 virtual machines across 14 systems that the company needs to release a batch, pay suppliers and staff, and meet regulators. Everything else recovers later from normal backups once the core is clean.

ItemFigureBasis
Protected scope~180 TB across 14 systemsTier 0 and tier 1 only, from business impact analysis
Nightly change sent to vault~2.4 TB~4% daily change = 7.2 TB, deduplicated about 3:1
Replication window2 hours a night on 10G2.4 TB at ~1 GB/s effective is ~40 minutes; 3x headroom for month end
Vault capacity~500 TB usable180 TB base + 90 days x 2.4 TB = ~400 TB, plus 25% growth
Scanning throughput~2.4 TB a night, full rescan weeklyChanged data scanned nightly; full copy of each system scanned every 7 days
Clean-room compute8 hosts, ~1,000 vCPU, 6 TB memory120 VMs at measured production size, no overcommit on memory
Clean-room storage~200 TB flashRestore at ~2 GB/s puts 180 TB back in about a day, tier 0 in hours

Vault, scanning, clean room and network are estimated at ₹12 to 16 crore, with annual support and testing effort on top. Recovery times in the outcomes are confirmed by the first full recovery test, not assumed.

How it is delivered

Five stages, ending in a real recovery test

The vault is itself a system that GxP data passes through, so it is designed, qualified and tested with quality assurance from the start, not after.

1

Assess and scope

Weeks 1 to 6

Business impact analysis, dependency maps for the 14 core systems, recovery order and targets agreed with the business and quality assurance.

Gate: Risk committee approves scope, recovery targets and budget.

2

Build vault and clean room

Weeks 5 to 16

Room, network, vault storage, scanning, vault control and clean-room hosts built with separate identity and keys. Gold images created.

Gate: Independent penetration test finds no path from production into the vault.

3

Qualify

Weeks 14 to 20

Installation and operational qualification of the vault and clean room under the company’s computerised system validation procedure.

Gate: Quality assurance signs the qualification report.

4

First full recovery test

Weeks 20 to 24

AD, ERP, LIMS and QMS rebuilt in the clean room from vault copies, then checked by application owners and QA. Every step timed.

Gate: Tier 0 in under 24 hours, batch-release systems in under 72 hours.

5

Operate and rehearse

Ongoing

Daily windows and scans, monthly restore of one system, quarterly full test of a recovery scenario, yearly test with the board’s crisis team.

Gate: Each test report reviewed by the risk committee.

Way back: Building the vault changes nothing in production: existing backups and the second site keep running as they are. If the vault or a scan job misbehaves, the window simply stays closed and the previous clean restore points remain locked and available.
Risks, handled up front

What could defeat a vault, and how the design closes each gap

RiskWhat could happenHow the design handles it
Dormant malware in backupsAttackers sit quietly for weeks so infected data is copied into the vaultScanning on every copy, 90-day retention so restore points older than the intrusion exist, and AD checks for changes that should not be there.
Insider or stolen vault credentialsSomeone shortens retention or deletes copiesCompliance-mode lock that no one can override, separate admin accounts with hardware tokens, two-person rule for any vault change.
Window left openThe link stays up and becomes a path inPorts disabled by default; an alarm and automatic shutdown if the link is up outside the window.
Recovery takes longer than plannedUntested dependencies slow the rebuildDependency-ordered runbooks, timed every quarter, with the slowest steps automated first.
Restored data not trusted by QABatch release cannot restart even when systems are backValidation checks agreed in advance, rehearsed with QA in every test, and audit trail continuity checked automatically.
What was optimised

What was optimised

22 h

Closed by default

The link is physically off for most of the day, and only the vault can open it.

14

Systems in scope

Only the minimum working company goes into the vault, which keeps cost and recovery time down.

3:1

Deduplicated copies

Only changed, unique data crosses each night, so a 10G link and a short window are enough.

Every copy

Scanned before trusted

A clean restore point is a measured fact, not a guess made during a crisis.

0

Shared credentials

Separate directory, keys and time source, so compromising production gives no way in.

Quarterly

Tested recovery

Recovery times come from real rebuilds, with QA involved every time.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
Copy an attacker cannot reachNoneVault copies with no shared admin path, locked for 90 days
Tier-0 recovery (AD, DNS, PKI)Unknown, never testedUnder 24 hours
Batch-release systems backWeeks, by estimate48 to 72 hours, validated
Knowing which copy is cleanGuesswork during an incidentRecorded for each system every night
Data loss for core systemsUp to 24 hours, if backups survivedUp to 24 hours, guaranteed copy available
Evidence for board and regulatorsBackup job reportsQuarterly recovery test reports with timings and QA sign-off

Targets are set in the scoping stage and confirmed by the first full recovery test. Recovery times depend on the size of the core scope and on how much of the runbook is automated.

Skills this draws on

What a team needs to deliver this

Cyber recovery architecture

Vault isolation, timed air gaps, retention locks and separate identity and key domains.

Backup and storage design

Deduplication, replication windows, capacity planning and restore throughput.

Active Directory recovery

Forest recovery, tier-0 hardening and credential resets in an isolated environment.

Malware analysis of backups

Scanning, entropy analysis and choosing clean restore points with evidence.

GxP computerised system validation

Qualification of recovery infrastructure and data integrity checks QA can sign.

Recovery testing and runbooks

Dependency-ordered runbooks, automation and quarterly tests that produce board-ready evidence.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

Could you rebuild if every backup you can reach today were encrypted?

Tell us which systems keep your business running and how your backups are managed today. We will come back with a plain view of where an attacker could reach your last copy, what a vault and clean room would need, and how quickly you could realistically recover.