Home / Deployment scenarios / Public cloud to private cloud
Reference deployment Cloud repatriation

Moving a SaaS platform from public cloud to its own private cloud

A payroll and HR software company had grown up on Azure. Six years later the bill had tripled, most of it for workloads that never changed size. This is how the move to a private cloud is designed so customers notice nothing except faster pages, and the company keeps the public cloud for the few things it is genuinely good at.

SectorB2B SaaS
Estate260 VMs, 2 AKS clusters
Programme7 months, 5 waves
ModelHybrid: private + burst
The situation

A cloud bill that grew faster than the business

The company runs payroll and HR software for about 1,200 mid-size and large employers in India. Everything lives on Azure: around 260 virtual machines, two Kubernetes clusters, managed PostgreSQL and SQL, and about 180 TB of documents in Blob storage.

Monthly spend had climbed from roughly ₹40 lakh to ₹1.4 crore in three years. When the finance team broke it down, about 80% of the cost came from workloads that run at the same size all year: the core application, databases and document store. Data leaving the cloud to customers, around 60 TB a month, had become a line item on its own.

At the same time, larger customers, including two banks, were asking harder questions in audits about where payroll data physically sits and who can reach it. The board asked a simple question: is there a cheaper way to run the same thing, without risking a single payroll run?

What could not be compromised

  • No payroll run can be put at risk. Nothing moves between the 25th and the 5th of any month.
  • No customer-visible downtime longer than 30 minutes for any service, with a tested way back.
  • ISO 27001 and SOC 2 controls must stay intact through the move, with evidence for auditors.
  • A platform team of nine engineers who know Azure well, but have not run OpenStack or Ceph.
  • Tax season brings a 3x peak for about six weeks a year. That capacity cannot sit idle all year.
Options weighed

Four ways forward, compared honestly

Repatriation is not automatically the right answer. Each option was costed over five years on the same workload data.

OptionWhat worksWhat does notVerdict
Stay and optimiseReserved instances and rightsizing cut 25 to 30%. No migration risk.Egress and managed-service premiums remain. Savings plateau after year one.Do first, regardless
Move to another public cloudPossible discounts in year one.Same cost structure, a full migration for a one-off saving.Rejected
Full repatriationLowest steady-state cost.Tax-season peak needs 3x hardware sitting idle 46 weeks a year.Rejected
Hybrid: steady state private, burst and archive on publicLowest five-year cost with peak covered. Data location becomes simple to evidence.Two environments to run, so tooling must be shared from day one.Chosen
Target architecture

Private cloud for the steady 80%, public cloud for the peaks

Production runs in a colocation data centre on an open-source stack. A second site holds the database replica, immutable backups and everything needed to rebuild. Azure stays connected through a private link for burst capacity and cold archive.

Scroll sideways to see the whole diagram →
INTERNET AND CUSTOMERSPUBLIC CLOUD (KEPT)PRIVATE CLOUD, COLOCATION DC 1DC 2: RECOVERYCustomer users1,200 companies, HR + payrollCDN and WAFTLS, bot and DDoS filteringPartner APIsbanks, tax, attendanceBurst node poolscales out for tax seasonExit estate260 VMs + managed SQLMOVINGCold archive tier7-year payroll recordsLoad balancersHA pair, L4 + L7Edge firewallsactive-standby pairKubernetes: production3 masters, 40 workersKubernetes: non-proddev, QA, stagingPostgreSQL HAPatroni, 3 nodes, NVMeOpenStack + KVM24 hosts, ~4,600 vCPUObservabilitymetrics, logs, tracesCeph block + object10 nodes, ~920 TB rawSpine-leaf network fabric2 x 100G spine, 25G to every hostDB standbyasync, ~5 min lagImmutable backupobject lock, 35 daysRecovery runbooksrebuild in 1 hourHTTPSSQLvolumes10G link1wave replication2replica3nightlytiering after 13 monthsUser or API trafficData / replicationControl / API callScheduled copy
Numbered flows: (1) workloads move wave by wave while data replicates from the old estate, (2) database replica to the recovery site, (3) nightly immutable backups. Burst traffic uses the private interconnect, never the internet.
Building blockWhy it is there
1 Edge firewalls and load balancersOne controlled entry point for customers and partners. Behind the same CDN and WAF the company already uses, so DNS cut-over is just a change of origin.
2 Kubernetes on OpenStackProduction and non-production clusters on KVM virtual machines. Same container images and Helm charts as on Azure, so applications do not change.
3 PostgreSQL with PatroniThree-node HA cluster on local NVMe replaces the managed database. Automatic failover inside the site, asynchronous replica in the second site.
4 Ceph block and object storageVolumes for VMs and containers, and an S3-compatible store for 180 TB of payslips and documents. Erasure coding keeps the object tier efficient.
5 Spine-leaf fabricTwo 100G spines and 25G to every host. Storage and application traffic on separate VLANs so a rebuild never slows payroll.
6 Burst pool and archive on AzureAn AKS node pool that scales out only in tax season, reached over a 2 x 10G private interconnect. Records older than 13 months tier to Azure cold archive.
7 Recovery siteDatabase replica, immutable backups with 35-day object lock, and runbooks to rebuild the platform from code. Tested every quarter.
Sizing, worked out

From 90 days of real usage, not from the current VM list

Azure VMs averaged 22% CPU use across the estate. Sizing used the 95th percentile of 90 days of metrics, then added headroom for two years of growth and one host failure per rack.

ResourceUsed today (p95)Private cloud designHow it was arrived at
vCPU~2,300 in use (of 6,800 provisioned)24 hosts, 2 x 48 cores each, ~4,600 vCPU at 2:1p95 use, +40% growth, N+2 hosts
Memory~9 TB in use24 x 768 GB = 18 TBMemory is not overcommitted
Block storage~120 TB3-way replicated on NVMeDatabases and volumes
Object storage~180 TB, growing 4 TB/monthErasure coded 4+210 Ceph nodes, ~920 TB raw, ~60% used at go-live
Network60 TB/month egress2 x 10G IP transit, flat ratePeak 3.2 Gbps measured
Burst3x peak for 6 weeksAKS node pool, 0 to 30 nodesPay only for the weeks used

Hardware is sized for two years. A third rack position is reserved in the colocation contract so year-three growth is a purchase order, not a redesign.

How it is delivered

Five waves, each one reversible

Every wave runs in both places for two weeks before Azure is switched off. Nothing moves in the payroll window.

1

Discover and order

Weeks 1 to 4

Dependency map of all 260 VMs and 140 services, from traffic data. Hardware ordered in week 2 to cover lead times.

Gate: Every service has an owner, a wave and a rollback plan.

2

Build the platform

Weeks 5 to 14

Racks, fabric, OpenStack, Ceph and Kubernetes built from code. Monitoring and backup working before any workload lands.

Gate: Failure tests passed: host, switch, disk and site link.

3

Non-production first

Weeks 15 to 18

Dev, QA and staging move. The platform team runs it daily and fixes what is awkward.

Gate: Two weeks of normal development with no platform tickets open.

4

Production waves

Weeks 19 to 28

Four waves, smallest risk first. Databases replicate ahead of each wave, DNS weights shift 10%, 50%, 100%.

Gate: Error rate and latency equal or better for 14 days.

5

Decommission

Weeks 29 to 30

Azure resources removed except burst pool and archive. Contracts and reservations closed.

Gate: Bill reconciled, audit evidence pack signed off.

Way back: During each wave, traffic weights can move back to Azure in minutes and the database replicates in both directions until sign-off. No wave closes until its rollback has been rehearsed.
Risks, handled up front

What could go wrong, and what is already in the plan

RiskWhat could happenHow the design handles it
Team skillsNine engineers inherit OpenStack and Ceph overnightHands-on training from week 5, runbooks written with the team, and six months of co-run support after go-live.
Hidden managed-service dependenciesAn application quietly relies on an Azure-only serviceFound in discovery from traffic and code scans. Service Bus and Key Vault mapped to RabbitMQ and Vault before wave 1.
Hardware lead timeServers or switches arrive late and push the programme into payrollOrdered in week 2 with dates in the contract. Waves are planned around freeze windows with slack.
Database cut-overData lost or out of sync at switch-overLogical replication for days ahead of each wave, checksums compared, cut-over rehearsed twice on a copy.
Single colocation siteFacility issue takes the platform downSecond site holds a replica and backups. Rebuild runbook tested every quarter with a measured time.
What was optimised

Where the savings actually come from

2:1

Right-sized compute

Sized from real p95 usage, not from 6,800 provisioned vCPU. Overcommit only on CPU, never on memory.

1.5x

Efficient object storage

Erasure coding 4+2 instead of three full copies for 180 TB of documents.

Flat

Predictable egress

Customer downloads move from per-GB billing to flat-rate IP transit in the colocation facility.

0

Licence lock-in

OpenStack, KVM, Ceph, Kubernetes and PostgreSQL are open source. Support is bought, not licences per core.

6 weeks

Peak paid only when used

Tax-season capacity stays on public cloud and scales to zero afterwards.

1 codebase

Same tooling both sides

Terraform, Helm and one monitoring stack cover private and public, so the team does not run two worlds.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
Monthly run cost~₹1.4 crore45 to 55% lower from year two, including hardware over five years, colocation and support
Payback on hardware-14 to 18 months
Egress cost60 TB/month at per-GB ratesFixed, inside the IP transit contract
Page load for Indian users (p95)~620 msUnder 400 ms, with the platform closer to users
RecoveryDepends on cloud regionRPO about 5 minutes, RTO 1 hour, tested quarterly
Data location evidenceShared responsibility documentsOwn racks, own access logs, one audit pack

Targets are set during design and measured after each wave. Cost figures depend on hardware pricing, colocation rates and growth, and are re-run with your own data during assessment.

Skills this draws on

What a team needs to deliver this

FinOps and cost modelling

Five-year cost models that compare options on the same usage data, including people and support.

OpenStack and KVM

Private cloud design, high availability for the control plane, and day-2 upgrades.

Ceph storage

Block and object design, erasure coding, failure domains and rebuild behaviour.

Kubernetes platforms

Cluster design, GitOps delivery and running the same workloads across private and public.

Database migration

PostgreSQL HA, logical replication and rehearsed cut-overs with checksums.

Migration programme design

Dependency mapping, wave planning around business calendars, and rollback at every step.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

Wondering if your cloud bill would look different on your own platform?

Send us a cost export and a rough list of what runs where. We will come back with a plain first view: which workloads are worth moving, which should stay, and what the five-year numbers look like.