Moving a SaaS platform from public cloud to its own private cloud
A payroll and HR software company had grown up on Azure. Six years later the bill had tripled, most of it for workloads that never changed size. This is how the move to a private cloud is designed so customers notice nothing except faster pages, and the company keeps the public cloud for the few things it is genuinely good at.
A cloud bill that grew faster than the business
The company runs payroll and HR software for about 1,200 mid-size and large employers in India. Everything lives on Azure: around 260 virtual machines, two Kubernetes clusters, managed PostgreSQL and SQL, and about 180 TB of documents in Blob storage.
Monthly spend had climbed from roughly ₹40 lakh to ₹1.4 crore in three years. When the finance team broke it down, about 80% of the cost came from workloads that run at the same size all year: the core application, databases and document store. Data leaving the cloud to customers, around 60 TB a month, had become a line item on its own.
At the same time, larger customers, including two banks, were asking harder questions in audits about where payroll data physically sits and who can reach it. The board asked a simple question: is there a cheaper way to run the same thing, without risking a single payroll run?
What could not be compromised
- No payroll run can be put at risk. Nothing moves between the 25th and the 5th of any month.
- No customer-visible downtime longer than 30 minutes for any service, with a tested way back.
- ISO 27001 and SOC 2 controls must stay intact through the move, with evidence for auditors.
- A platform team of nine engineers who know Azure well, but have not run OpenStack or Ceph.
- Tax season brings a 3x peak for about six weeks a year. That capacity cannot sit idle all year.
Four ways forward, compared honestly
Repatriation is not automatically the right answer. Each option was costed over five years on the same workload data.
Private cloud for the steady 80%, public cloud for the peaks
Production runs in a colocation data centre on an open-source stack. A second site holds the database replica, immutable backups and everything needed to rebuild. Azure stays connected through a private link for burst capacity and cold archive.
| Building block | Why it is there |
|---|---|
| 1 Edge firewalls and load balancers | One controlled entry point for customers and partners. Behind the same CDN and WAF the company already uses, so DNS cut-over is just a change of origin. |
| 2 Kubernetes on OpenStack | Production and non-production clusters on KVM virtual machines. Same container images and Helm charts as on Azure, so applications do not change. |
| 3 PostgreSQL with Patroni | Three-node HA cluster on local NVMe replaces the managed database. Automatic failover inside the site, asynchronous replica in the second site. |
| 4 Ceph block and object storage | Volumes for VMs and containers, and an S3-compatible store for 180 TB of payslips and documents. Erasure coding keeps the object tier efficient. |
| 5 Spine-leaf fabric | Two 100G spines and 25G to every host. Storage and application traffic on separate VLANs so a rebuild never slows payroll. |
| 6 Burst pool and archive on Azure | An AKS node pool that scales out only in tax season, reached over a 2 x 10G private interconnect. Records older than 13 months tier to Azure cold archive. |
| 7 Recovery site | Database replica, immutable backups with 35-day object lock, and runbooks to rebuild the platform from code. Tested every quarter. |
From 90 days of real usage, not from the current VM list
Azure VMs averaged 22% CPU use across the estate. Sizing used the 95th percentile of 90 days of metrics, then added headroom for two years of growth and one host failure per rack.
| Resource | Used today (p95) | Private cloud design | How it was arrived at |
|---|---|---|---|
| vCPU | ~2,300 in use (of 6,800 provisioned) | 24 hosts, 2 x 48 cores each, ~4,600 vCPU at 2:1 | p95 use, +40% growth, N+2 hosts |
| Memory | ~9 TB in use | 24 x 768 GB = 18 TB | Memory is not overcommitted |
| Block storage | ~120 TB | 3-way replicated on NVMe | Databases and volumes |
| Object storage | ~180 TB, growing 4 TB/month | Erasure coded 4+2 | 10 Ceph nodes, ~920 TB raw, ~60% used at go-live |
| Network | 60 TB/month egress | 2 x 10G IP transit, flat rate | Peak 3.2 Gbps measured |
| Burst | 3x peak for 6 weeks | AKS node pool, 0 to 30 nodes | Pay only for the weeks used |
Hardware is sized for two years. A third rack position is reserved in the colocation contract so year-three growth is a purchase order, not a redesign.
Five waves, each one reversible
Every wave runs in both places for two weeks before Azure is switched off. Nothing moves in the payroll window.
Discover and order
Weeks 1 to 4
Dependency map of all 260 VMs and 140 services, from traffic data. Hardware ordered in week 2 to cover lead times.
Gate: Every service has an owner, a wave and a rollback plan.
Build the platform
Weeks 5 to 14
Racks, fabric, OpenStack, Ceph and Kubernetes built from code. Monitoring and backup working before any workload lands.
Gate: Failure tests passed: host, switch, disk and site link.
Non-production first
Weeks 15 to 18
Dev, QA and staging move. The platform team runs it daily and fixes what is awkward.
Gate: Two weeks of normal development with no platform tickets open.
Production waves
Weeks 19 to 28
Four waves, smallest risk first. Databases replicate ahead of each wave, DNS weights shift 10%, 50%, 100%.
Gate: Error rate and latency equal or better for 14 days.
Decommission
Weeks 29 to 30
Azure resources removed except burst pool and archive. Contracts and reservations closed.
Gate: Bill reconciled, audit evidence pack signed off.
What could go wrong, and what is already in the plan
| Risk | What could happen | How the design handles it |
|---|---|---|
| Team skills | Nine engineers inherit OpenStack and Ceph overnight | Hands-on training from week 5, runbooks written with the team, and six months of co-run support after go-live. |
| Hidden managed-service dependencies | An application quietly relies on an Azure-only service | Found in discovery from traffic and code scans. Service Bus and Key Vault mapped to RabbitMQ and Vault before wave 1. |
| Hardware lead time | Servers or switches arrive late and push the programme into payroll | Ordered in week 2 with dates in the contract. Waves are planned around freeze windows with slack. |
| Database cut-over | Data lost or out of sync at switch-over | Logical replication for days ahead of each wave, checksums compared, cut-over rehearsed twice on a copy. |
| Single colocation site | Facility issue takes the platform down | Second site holds a replica and backups. Rebuild runbook tested every quarter with a measured time. |
Where the savings actually come from
Right-sized compute
Sized from real p95 usage, not from 6,800 provisioned vCPU. Overcommit only on CPU, never on memory.
Efficient object storage
Erasure coding 4+2 instead of three full copies for 180 TB of documents.
Predictable egress
Customer downloads move from per-GB billing to flat-rate IP transit in the colocation facility.
Licence lock-in
OpenStack, KVM, Ceph, Kubernetes and PostgreSQL are open source. Support is bought, not licences per core.
Peak paid only when used
Tax-season capacity stays on public cloud and scales to zero afterwards.
Same tooling both sides
Terraform, Helm and one monitoring stack cover private and public, so the team does not run two worlds.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Monthly run cost | ~₹1.4 crore | 45 to 55% lower from year two, including hardware over five years, colocation and support |
| Payback on hardware | - | 14 to 18 months |
| Egress cost | 60 TB/month at per-GB rates | Fixed, inside the IP transit contract |
| Page load for Indian users (p95) | ~620 ms | Under 400 ms, with the platform closer to users |
| Recovery | Depends on cloud region | RPO about 5 minutes, RTO 1 hour, tested quarterly |
| Data location evidence | Shared responsibility documents | Own racks, own access logs, one audit pack |
Targets are set during design and measured after each wave. Cost figures depend on hardware pricing, colocation rates and growth, and are re-run with your own data during assessment.
What a team needs to deliver this
FinOps and cost modelling
Five-year cost models that compare options on the same usage data, including people and support.
OpenStack and KVM
Private cloud design, high availability for the control plane, and day-2 upgrades.
Ceph storage
Block and object design, erasure coding, failure domains and rebuild behaviour.
Kubernetes platforms
Cluster design, GitOps delivery and running the same workloads across private and public.
Database migration
PostgreSQL HA, logical replication and rehearsed cut-overs with checksums.
Migration programme design
Dependency mapping, wave planning around business calendars, and rollback at every step.
Wondering if your cloud bill would look different on your own platform?
Send us a cost export and a rough list of what runs where. We will come back with a plain first view: which workloads are worth moving, which should stay, and what the five-year numbers look like.