Home / Deployment scenarios / Multi-tenant DR as a Service
Reference deployment Disaster recovery

Disaster recovery as a service for 50 customers, without any of them seeing each other

A regional data-centre operator already hosts racks for many mid-size companies. Most of them have no real DR: a backup on tape or in a second room down the corridor. This design turns spare capacity into a shared DR service, where each customer replicates its servers to the operator, tests recovery on its own whenever it likes, and pays per protected VM, while staying fully isolated from every other tenant.

SectorData-centre operator
Customers40 to 60, 25 to 150 VMs each
Programme8 months to first 10 tenants
ModelShared platform, isolated tenants
The situation

Customers asking for DR, and no product to sell them

The operator runs a Tier III facility serving manufacturers, hospitals, NBFCs, schools and regional retailers. Customers increasingly ask the same thing: auditors, insurers and lenders want to see a DR plan with a tested recovery time, and the customers cannot justify a second data centre of their own.

Today the operator can only offer rack space in a second hall and leave the rest to the customer. A few customers buy DR from a hyperscaler, which means a different contract, data outside the region and a skills gap when it matters. Others simply hope. Sales has a list of over 40 prospects with between 25 and 150 VMs each, running a mix of VMware, Hyper-V and KVM.

Leadership wants a standard product: replicate a customer’s VMs into the operator’s cloud, let the customer test and fail over from a portal, guarantee that no tenant can ever see another tenant’s data or network, and bill by protected VM so the price is easy to understand. The platform has to pay for itself well before the hardware is old.

What could not be compromised

  • Strict tenant isolation: separate networks, storage pools, encryption keys and portal roles. One tenant’s test or failover can never affect another.
  • Customers run VMware, Hyper-V and KVM at their own sites. The service cannot require them to change hypervisor.
  • Default targets of RPO 15 minutes and RTO 4 hours for the standard tier, written into the service contract.
  • Customers must be able to run their own DR tests without raising a ticket, and get a report they can hand to an auditor.
  • Capital outlay has to be recovered within about 30 months at realistic take-up, with clear pricing per protected VM.
Options weighed

Four ways to offer DR, compared honestly

Each option was modelled for 50 tenants averaging 70 VMs, on margin, isolation, recovery targets and the operator’s ability to support it with a team of eight.

OptionWhat worksWhat does notVerdict
Backup-based recovery onlyCheapest. Uses backup software the operator already sells.Restores take a day or more for a whole site, and RPO is the last nightly backup.Kept as a lower tier
Resell a hyperscaler DR serviceNo platform to build.Thin margin, data leaves the region, and the operator owns the customer relationship but not the recovery.Rejected
Dedicated DR stack per customerStrongest isolation, simple to explain.Hardware, licences and support per customer make it uneconomic below about 300 VMs.Rejected
Shared platform with isolated tenants and self-serviceBest economics at 40 to 60 tenants, one team runs it, isolation enforced in network, storage and portal.Isolation must be designed and tested with care. Needs a portal and billing that work from day one.Chosen
Target architecture

One shared platform, a separate world for every tenant

Compute, storage and orchestration are shared and run by the operator. Everything a tenant touches is separated: each tenant arrives over its own tunnel into its own VRF, its replicas live in its own storage pool under its own key, and its recovered VMs boot into its own network. Tests boot a copy into a bubble with no route out, so a test never collides with production.

Scroll sideways to see the whole diagram →
CUSTOMER SITESOPERATOR DR CLOUD: SHARED PLATFORMISOLATED TENANT ZONESCustomer IT adminsone login per tenant, MFATenant A siteVMware, 80 VMsTenant B siteHyper-V, 45 VMsTenant C siteKVM, 120 VMsSelf-service portaltenant-scoped roles, auditTenant edge gatewaysIPsec or MPLS, VRF per tenantReplication gatewayscompress, encrypt, journalOperator NOCRPO lag, health, ticketsDR orchestratorrunbooks, boot orderKVM recovery poolOpenStack, 24 hostsReplica storageCeph, pool per tenantUsage and billingper protected VM, monthlyTenant A networkown VRF, VXLAN and firewallTenant B networkown VRF, VXLAN and firewallTenant C networkown VRF, VXLAN and firewallTest bubblecopy of tenant, no route outTESTown tunnel4boot planvolumesRPO lagusage12356User or API trafficData / replicationControl / API callLogging / management
Numbered flows: (1) each tenant replicates over its own link into its own VRF, (2) changed blocks are written to a journal in that tenant’s storage pool, (3) the customer starts a test or failover from the portal, (4) the orchestrator boots VMs in runbook order on the shared KVM pool, (5) on failover, VMs come up in the tenant’s isolated network, (6) on a test, a copy boots into a test bubble with no route to anything else.
Building blockWhy it is there
1 Replication at customer sitesA small replication appliance per site, hypervisor-level for VMware and Hyper-V and agent-based for KVM and physical servers. It sends changed blocks continuously, compressed and encrypted, with a bandwidth cap set per customer.
2 Tenant edge gatewaysA pair of edge routers and firewalls terminate each tenant’s IPsec tunnel or MPLS link into a dedicated VRF. Tenant address ranges may overlap, and nothing routes between VRFs.
3 Replication gateways and journalsReceive each tenant’s stream, write it to a journal and keep recovery points every few minutes for 72 hours, plus daily points for 7 days. Lag per VM is reported to the NOC and to the customer.
4 Replica storage on CephOne storage pool and one encryption key per tenant. Replica disks sit on an erasure-coded pool; journals sit on a replicated NVMe pool so recovery points can be mounted quickly.
5 KVM recovery pool on OpenStackShared compute that sits mostly idle, running only tests and real failovers. Each tenant is an OpenStack project with quotas, so one tenant cannot consume another’s reserved capacity.
6 Tenant networks and test bubblesEVPN-VXLAN segments per tenant, each with its own virtual firewall and the same IP addresses as the customer’s site. A test bubble is a further isolated segment with no gateway at all.
7 Portal and orchestratorCustomers see only their own VMs, set boot order and IP mapping, run tests and download a report. The orchestrator executes runbooks: boot databases first, wait for health checks, then applications, then re-point DNS.
8 Metering and billingCounts protected VMs, protected storage above the included allowance, and recovery hours used outside tests. Feeds the operator’s billing system every month.
Sizing, worked out

Sized for 3,500 protected VMs, not for every tenant failing at once

The design assumes 50 tenants at about 70 VMs each, with around 250 GB used per VM and about 3% of data changing each day. Recovery capacity is sized for the realistic worst case: a few large tenants failing over at once while others run tests.

ItemFigureBasis
Protected VMs~3,500, designed for 4,00050 tenants x ~70 VMs, plus 15% headroom
Replica capacity~875 TB3,500 VMs x ~250 GB used
Daily change and journal~26 TB a day, ~80 TB journal3,500 x 250 GB x 3% change; 72 hours of recovery points
Storage~1 PB usable, ~1.6 PB rawReplicas on 4+2 erasure coding (1.5x), journals on 3x replicated NVMe
Recovery compute24 hosts, ~4,600 vCPU, 18 TB RAM30% of VMs running at once: ~1,050 VMs x 4 vCPU at 2:1, 16 GB each
Customer link~25 Mbps average, ~75 Mbps peak70 VMs x 7.5 GB/day = 525 GB/day, compressed 2:1, 3x peak factor
Operator ingress2 x 10G, plus MPLS handoffs50 tenants x ~25 Mbps average = ~1.25 Gbps, peaks near 4 Gbps

Compute is the cost that scales with risk, not with tenants. Contracts offer reserved recovery capacity for customers who need guaranteed resources during a regional event, and standard capacity on a shared basis for the rest. A third compute rack is pre-wired so reserved capacity can be added as tenants buy it.

How it is delivered

Build, prove with three anchor tenants, then scale

The platform does not go on sale until three real customers have failed over and back, and an outside tester has failed to break isolation.

1

Design and commercial model

Weeks 1 to 5

Service tiers, pricing per protected VM, isolation model, replication approach for each hypervisor. Hardware ordered in week 4.

Gate: Pricing model approved, isolation design reviewed by an independent security assessor.

2

Build the platform

Weeks 6 to 16

Edge, Ceph, OpenStack, orchestrator and portal built from code. Two internal “tenants” created to test every flow.

Gate: Isolation penetration test passed: no route, storage or portal access between tenants.

3

Anchor tenants

Weeks 17 to 24

Three friendly customers, one per hypervisor, replicate production. Each runs a test and a planned full failover and failback.

Gate: All three meet RPO 15 minutes and RTO 4 hours in measured drills.

4

Productise

Weeks 25 to 30

Runbook templates, onboarding checklist, billing integration, customer-facing DR test reports and NOC procedures.

Gate: A new tenant onboarded by the operations team alone in under 10 working days.

5

Scale

Month 8 onwards

Onboard in waves of five tenants a month, with capacity reviews every quarter.

Gate: Storage and compute use reviewed against reservations before each wave.

Way back: Replication is non-disruptive at the customer site: if a tenant leaves or onboarding stalls, the appliance is removed and production is untouched. Any failover can be failed back once the customer site is healthy, using reverse replication of changes made while running at the operator.
Risks, handled up front

What could go wrong, and what is already in the plan

RiskWhat could happenHow the design handles it
Isolation failureA misconfiguration lets one tenant reach another’s network or dataVRF, storage pool, encryption key and portal role per tenant, all created from one template. Independent penetration test before launch and every year.
Regional disaster hits many tenantsA flood or grid failure takes out several customers in the same city at onceReserved capacity tiers in contracts, standard tier served on a fair-share basis, and a clear priority order agreed in advance.
Replication lagThin customer links mean RPO is quietly missedBandwidth checked before onboarding, lag alerts per VM to the NOC and the customer, and the SLA tied to measured link capacity.
Untested runbooksA failover works for VMs but the application does not start in the right orderEvery tenant must pass a test failover before go-live and at least twice a year, with a report signed by the customer.
Slow take-upFewer tenants than planned and the platform does not pay backHardware bought for the first 25 tenants only, with racks pre-wired for growth. Backup-based DR offered as an entry tier to bring customers in.
What was optimised

What makes the economics work

1.5x

Erasure-coded replicas

Replica disks sit on 4+2 erasure coding, so a petabyte of protected data needs about 1.6 PB raw, not 3 PB.

30%

Compute sized for reality

Recovery hosts are sized for the worst likely day, not for every tenant failing at once.

0

Hypervisor changes

Customers keep VMware, Hyper-V or KVM at their own site. Everything lands on KVM at the operator.

1 template

Tenants built from code

Each tenant’s VRF, pool, key, quota and portal role come from one template, so isolation does not depend on someone’s memory.

Self-service

Tests without tickets

Customers run their own tests whenever they like. The NOC only steps in when something fails.

Per VM

Simple pricing

One price per protected VM with storage included up to an allowance, so customers can budget and sales can quote in minutes.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
Recovery point (standard tier)Last backup, often 24 hours15 minutes, typically under 5
Recovery time (standard tier)Days, if at all4 hours for a whole tenant, measured in drills
DR test effort for a customerA weekend and a project planSelf-service, a few hours, report included
Price to customerSecond DC or nothing₹1,800 to ₹2,800 per protected VM per month
Platform investment-₹9 to 11 crore for the first 4,000 VMs
Payback-24 to 30 months at 35 to 50 tenants

Commercial figures depend on hardware pricing, replication licence costs, take-up rate and the mix of reserved and standard capacity. They are rebuilt with the operator’s own costs and pipeline during design.

Skills this draws on

What a team needs to deliver this

DR architecture

RPO and RTO design, replication approaches across hypervisors, and failback that actually works.

Multi-tenant networking

VRF and EVPN-VXLAN design, overlapping address ranges and test bubbles that cannot leak.

OpenStack and KVM

Projects, quotas and capacity planning for compute that sits idle until it matters.

Ceph storage

Per-tenant pools, encryption keys, erasure coding and journal performance.

Runbook automation

Boot order, health checks, IP mapping and DNS changes turned into repeatable runbooks.

Service design and pricing

Turning a platform into a product: tiers, SLAs, metering and a cost model that pays back.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

Thinking of turning your data centre capacity into a DR service?

Share your facility, the hardware you already have and a rough list of prospects. We will come back with a plain view of the platform, the isolation model, the pricing per VM and when it pays for itself.