Disaster recovery as a service for 50 customers, without any of them seeing each other
A regional data-centre operator already hosts racks for many mid-size companies. Most of them have no real DR: a backup on tape or in a second room down the corridor. This design turns spare capacity into a shared DR service, where each customer replicates its servers to the operator, tests recovery on its own whenever it likes, and pays per protected VM, while staying fully isolated from every other tenant.
Customers asking for DR, and no product to sell them
The operator runs a Tier III facility serving manufacturers, hospitals, NBFCs, schools and regional retailers. Customers increasingly ask the same thing: auditors, insurers and lenders want to see a DR plan with a tested recovery time, and the customers cannot justify a second data centre of their own.
Today the operator can only offer rack space in a second hall and leave the rest to the customer. A few customers buy DR from a hyperscaler, which means a different contract, data outside the region and a skills gap when it matters. Others simply hope. Sales has a list of over 40 prospects with between 25 and 150 VMs each, running a mix of VMware, Hyper-V and KVM.
Leadership wants a standard product: replicate a customer’s VMs into the operator’s cloud, let the customer test and fail over from a portal, guarantee that no tenant can ever see another tenant’s data or network, and bill by protected VM so the price is easy to understand. The platform has to pay for itself well before the hardware is old.
What could not be compromised
- Strict tenant isolation: separate networks, storage pools, encryption keys and portal roles. One tenant’s test or failover can never affect another.
- Customers run VMware, Hyper-V and KVM at their own sites. The service cannot require them to change hypervisor.
- Default targets of RPO 15 minutes and RTO 4 hours for the standard tier, written into the service contract.
- Customers must be able to run their own DR tests without raising a ticket, and get a report they can hand to an auditor.
- Capital outlay has to be recovered within about 30 months at realistic take-up, with clear pricing per protected VM.
Four ways to offer DR, compared honestly
Each option was modelled for 50 tenants averaging 70 VMs, on margin, isolation, recovery targets and the operator’s ability to support it with a team of eight.
One shared platform, a separate world for every tenant
Compute, storage and orchestration are shared and run by the operator. Everything a tenant touches is separated: each tenant arrives over its own tunnel into its own VRF, its replicas live in its own storage pool under its own key, and its recovered VMs boot into its own network. Tests boot a copy into a bubble with no route out, so a test never collides with production.
| Building block | Why it is there |
|---|---|
| 1 Replication at customer sites | A small replication appliance per site, hypervisor-level for VMware and Hyper-V and agent-based for KVM and physical servers. It sends changed blocks continuously, compressed and encrypted, with a bandwidth cap set per customer. |
| 2 Tenant edge gateways | A pair of edge routers and firewalls terminate each tenant’s IPsec tunnel or MPLS link into a dedicated VRF. Tenant address ranges may overlap, and nothing routes between VRFs. |
| 3 Replication gateways and journals | Receive each tenant’s stream, write it to a journal and keep recovery points every few minutes for 72 hours, plus daily points for 7 days. Lag per VM is reported to the NOC and to the customer. |
| 4 Replica storage on Ceph | One storage pool and one encryption key per tenant. Replica disks sit on an erasure-coded pool; journals sit on a replicated NVMe pool so recovery points can be mounted quickly. |
| 5 KVM recovery pool on OpenStack | Shared compute that sits mostly idle, running only tests and real failovers. Each tenant is an OpenStack project with quotas, so one tenant cannot consume another’s reserved capacity. |
| 6 Tenant networks and test bubbles | EVPN-VXLAN segments per tenant, each with its own virtual firewall and the same IP addresses as the customer’s site. A test bubble is a further isolated segment with no gateway at all. |
| 7 Portal and orchestrator | Customers see only their own VMs, set boot order and IP mapping, run tests and download a report. The orchestrator executes runbooks: boot databases first, wait for health checks, then applications, then re-point DNS. |
| 8 Metering and billing | Counts protected VMs, protected storage above the included allowance, and recovery hours used outside tests. Feeds the operator’s billing system every month. |
Sized for 3,500 protected VMs, not for every tenant failing at once
The design assumes 50 tenants at about 70 VMs each, with around 250 GB used per VM and about 3% of data changing each day. Recovery capacity is sized for the realistic worst case: a few large tenants failing over at once while others run tests.
| Item | Figure | Basis |
|---|---|---|
| Protected VMs | ~3,500, designed for 4,000 | 50 tenants x ~70 VMs, plus 15% headroom |
| Replica capacity | ~875 TB | 3,500 VMs x ~250 GB used |
| Daily change and journal | ~26 TB a day, ~80 TB journal | 3,500 x 250 GB x 3% change; 72 hours of recovery points |
| Storage | ~1 PB usable, ~1.6 PB raw | Replicas on 4+2 erasure coding (1.5x), journals on 3x replicated NVMe |
| Recovery compute | 24 hosts, ~4,600 vCPU, 18 TB RAM | 30% of VMs running at once: ~1,050 VMs x 4 vCPU at 2:1, 16 GB each |
| Customer link | ~25 Mbps average, ~75 Mbps peak | 70 VMs x 7.5 GB/day = 525 GB/day, compressed 2:1, 3x peak factor |
| Operator ingress | 2 x 10G, plus MPLS handoffs | 50 tenants x ~25 Mbps average = ~1.25 Gbps, peaks near 4 Gbps |
Compute is the cost that scales with risk, not with tenants. Contracts offer reserved recovery capacity for customers who need guaranteed resources during a regional event, and standard capacity on a shared basis for the rest. A third compute rack is pre-wired so reserved capacity can be added as tenants buy it.
Build, prove with three anchor tenants, then scale
The platform does not go on sale until three real customers have failed over and back, and an outside tester has failed to break isolation.
Design and commercial model
Weeks 1 to 5
Service tiers, pricing per protected VM, isolation model, replication approach for each hypervisor. Hardware ordered in week 4.
Gate: Pricing model approved, isolation design reviewed by an independent security assessor.
Build the platform
Weeks 6 to 16
Edge, Ceph, OpenStack, orchestrator and portal built from code. Two internal “tenants” created to test every flow.
Gate: Isolation penetration test passed: no route, storage or portal access between tenants.
Anchor tenants
Weeks 17 to 24
Three friendly customers, one per hypervisor, replicate production. Each runs a test and a planned full failover and failback.
Gate: All three meet RPO 15 minutes and RTO 4 hours in measured drills.
Productise
Weeks 25 to 30
Runbook templates, onboarding checklist, billing integration, customer-facing DR test reports and NOC procedures.
Gate: A new tenant onboarded by the operations team alone in under 10 working days.
Scale
Month 8 onwards
Onboard in waves of five tenants a month, with capacity reviews every quarter.
Gate: Storage and compute use reviewed against reservations before each wave.
What could go wrong, and what is already in the plan
| Risk | What could happen | How the design handles it |
|---|---|---|
| Isolation failure | A misconfiguration lets one tenant reach another’s network or data | VRF, storage pool, encryption key and portal role per tenant, all created from one template. Independent penetration test before launch and every year. |
| Regional disaster hits many tenants | A flood or grid failure takes out several customers in the same city at once | Reserved capacity tiers in contracts, standard tier served on a fair-share basis, and a clear priority order agreed in advance. |
| Replication lag | Thin customer links mean RPO is quietly missed | Bandwidth checked before onboarding, lag alerts per VM to the NOC and the customer, and the SLA tied to measured link capacity. |
| Untested runbooks | A failover works for VMs but the application does not start in the right order | Every tenant must pass a test failover before go-live and at least twice a year, with a report signed by the customer. |
| Slow take-up | Fewer tenants than planned and the platform does not pay back | Hardware bought for the first 25 tenants only, with racks pre-wired for growth. Backup-based DR offered as an entry tier to bring customers in. |
What makes the economics work
Erasure-coded replicas
Replica disks sit on 4+2 erasure coding, so a petabyte of protected data needs about 1.6 PB raw, not 3 PB.
Compute sized for reality
Recovery hosts are sized for the worst likely day, not for every tenant failing at once.
Hypervisor changes
Customers keep VMware, Hyper-V or KVM at their own site. Everything lands on KVM at the operator.
Tenants built from code
Each tenant’s VRF, pool, key, quota and portal role come from one template, so isolation does not depend on someone’s memory.
Tests without tickets
Customers run their own tests whenever they like. The NOC only steps in when something fails.
Simple pricing
One price per protected VM with storage included up to an allowance, so customers can budget and sales can quote in minutes.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Recovery point (standard tier) | Last backup, often 24 hours | 15 minutes, typically under 5 |
| Recovery time (standard tier) | Days, if at all | 4 hours for a whole tenant, measured in drills |
| DR test effort for a customer | A weekend and a project plan | Self-service, a few hours, report included |
| Price to customer | Second DC or nothing | ₹1,800 to ₹2,800 per protected VM per month |
| Platform investment | - | ₹9 to 11 crore for the first 4,000 VMs |
| Payback | - | 24 to 30 months at 35 to 50 tenants |
Commercial figures depend on hardware pricing, replication licence costs, take-up rate and the mix of reserved and standard capacity. They are rebuilt with the operator’s own costs and pipeline during design.
What a team needs to deliver this
DR architecture
RPO and RTO design, replication approaches across hypervisors, and failback that actually works.
Multi-tenant networking
VRF and EVPN-VXLAN design, overlapping address ranges and test bubbles that cannot leak.
OpenStack and KVM
Projects, quotas and capacity planning for compute that sits idle until it matters.
Ceph storage
Per-tenant pools, encryption keys, erasure coding and journal performance.
Runbook automation
Boot order, health checks, IP mapping and DNS changes turned into repeatable runbooks.
Service design and pricing
Turning a platform into a product: tiers, SLAs, metering and a cost model that pays back.
Other scenarios
Three-site core banking DR
A small finance bank protects its core banking and payment switch across three sites, with zero data loss to a near DR site and a far site in another seismic zone.
Read →Cyber recoveryCyber recovery vault for ransomware
A listed pharmaceutical company builds an isolated recovery vault and clean room, so it can rebuild after ransomware from copies attackers cannot reach.
Read →Security operations and DRSOC infrastructure across DC and DR
A private bank builds its own Security Operations Centre across two data centres, sized from 60,000 events per second, with a DR SIEM that can take over.
Read →Thinking of turning your data centre capacity into a DR service?
Share your facility, the hardware you already have and a rough list of prospects. We will come back with a plain view of the platform, the isolation model, the pricing per VM and when it pays for itself.