Home / Disaster Recovery / DR in the public cloud
DR guide

Disaster recovery in the public cloud

Building a second data centre just for DR means buying hardware that sits idle most of the year. A public cloud region can act as the DR site instead: you pay to keep data copied there all the time, and pay for servers only when you test or actually fail over. For many organisations that changes DR from a capital project into a modest monthly cost. It also brings its own catches, which this guide covers.

Three ways to use the cloud for DR

SetupWhat it meansGood forWatch out for
On premises to cloudProduction stays in your data centre; a cloud region is the DR siteOrganisations without a second data centre, or replacing an ageing oneBandwidth for replication, licences at DR, and the cost of bringing data back
Cloud region to cloud regionProduction runs in one cloud region, DR in another region of the same providerWorkloads already in the cloudBoth regions share the same provider, account and identity system
Cloud to a different provider or back on premisesDR outside the provider that runs productionRegulated sectors worried about dependence on one providerTwo platforms to run and test, and services that do not translate one to one

The most common design: on premises to a cloud region

Scroll sideways to see the whole diagram →
YOUR DATA CENTRE (PRIMARY)CLOUD REGION IN INDIA (DR)Application VMsrunning, serving usersReplication agentblock-level, continuousDatabasesprimary, takes all writesBackupsdaily, kept on siteNetwork edgefirewalls, WAN, internetRecovery serversnot running until a drill or disasterOFFStaging storagelow-cost disks, always onDatabase replicasmall instance, always onBackup copyobject storage, lockedCloud networkpre-built, same rules as primary1continuous2log shipping3daily copyVPN or private link4Data / replicationScheduled copyUser or API trafficControl / API call
1. A replication agent copies each server’s disks continuously to low-cost staging storage in the cloud region. 2. Databases ship their logs to a small replica that is always running, so the RPO stays in minutes. 3. Backups are copied daily to object storage with a lock, so they survive ransomware at the primary site. 4. Recovery servers do not run day to day. At a drill or a disaster they are launched from the staging storage, which is why the standing cost is low.

A worked example

Forty servers, 20 TB of data, an RPO of 15 minutes and an RTO of four hours, with DR in an Indian cloud region. These are planning ranges to show where the money goes, not a quote. Prices vary by provider, region, commitment and discounts.

ItemBasisPlanning figure
Staging storage for replicated disks20 TB × ₹5–8 per GB per month₹1.0–1.6 lakh
DR tooling, per protected server40 servers × ₹1,500–2,500₹0.6–1.0 lakh
Small database replica, always onOne modest instance with storage₹0.4–0.8 lakh
Backup copy in locked object storage30 TB with retention × ₹1.5–2.5 per GB₹0.45–0.75 lakh
ConnectivitySite-to-site VPN or a small private link₹0.3–0.8 lakh
Standing cost per monthSum of the aboveabout ₹3–5 lakh
Compute during a drill or disaster40 servers × ₹300–600 per day₹12,000–24,000 per day
Failback data transfer20 TB × ₹7–9 per GB out of the cloud₹1.4–1.8 lakh, once

For comparison, a warm-standby second data centre for the same estate typically needs ₹1.5–3 crore of hardware up front, plus ₹3–6 lakh a month for colocation, power and connectivity. Cloud DR is usually cheaper for this size of estate. It becomes less attractive when DR servers must run all the time, or when large volumes of data move out of the cloud often.

The catches

Where the data may live

Some data must stay in India, for example payment data under RBI rules. Choose an Indian region (Mumbai, Hyderabad, Pune, Chennai or Delhi, depending on the provider) and check your sector regulator’s rules before you start.

Capacity on the day

In a regional event, many customers fail over at once. For critical tiers, reserve capacity or keep a small footprint always running rather than assuming servers will be available.

Licences at DR

Oracle, Microsoft SQL Server and Windows licences have specific rules for DR and for running in a public cloud. Check them during design, not after the first drill.

Getting data back

Data leaving the cloud is charged per GB. After a real failover, bringing tens of terabytes back to the primary site is a real cost and takes time. Plan failback as carefully as failover.

Bandwidth for replication

The link has to carry your daily change rate within the RPO. Measure the real change rate first; a guess here is the most common reason a cloud DR design misses its RPO.

One account, one key

If an attacker gets into the cloud account, both the DR servers and the backups are at risk. Keep backups in a separate account with locked retention and separate administrator credentials.

Common building blocks by provider

NeedAWSAzureOracle Cloud
Server replication and recoveryElastic Disaster RecoveryAzure Site RecoveryFull Stack Disaster Recovery
Locked backupsAWS Backup with Vault LockAzure Backup with immutable vaultsObject Storage retention rules
Private connectivityDirect Connect, Site-to-Site VPNExpressRoute, VPN GatewayFastConnect, Site-to-Site VPN
Indian regionsMumbai, HyderabadPune, Chennai, MumbaiMumbai, Hyderabad

Third-party tools such as Veeam and Zerto also replicate on-premises servers into these clouds, and are often the better choice when you want the same tool across several platforms.

Not sure where your DR stands?

Send us your application list and a short note on how backups work today. We will come back with a plain gap review: which systems are exposed, what a realistic recovery time looks like, and what it would take to close the gap.