Executive summary
Most organisations now run some workloads in a public cloud and some on infrastructure they control, and many are being pushed to rethink the second half by a VMware renewal. The question is no longer “cloud or not”. It is which workload runs where, at what five-year cost, and how to move without outages.
This paper covers placement, cost, a private cloud built on OpenStack, KVM and Ceph, a safe exit from VMware, and the landing zone and operating model that keep a hybrid estate under control.
Five things to take away
- Decide placement per workload, in writing. Load shape, data movement, regulation, managed services and team skills decide where it belongs. A blanket “cloud first” or “all private” policy is usually the expensive answer.
- Compare five-year cost, not first-year cost. Include people, data egress, licences, support and hardware refresh. Large, steady 24x7 estates often cost 25 to 40 percent less on a well-run private cloud; spiky or short-lived work is usually cheaper in public cloud.
- Private cloud means self-service, not just virtualisation. OpenStack on KVM with Ceph gives teams an API and portal on hardware you own, with no hypervisor licence. It needs a small skilled team and a firm upgrade habit.
- Leave VMware in waves, with a way back. Move dependency groups together, test side by side, and keep the source VM intact until the application owner signs off.
- Build the guardrails before the workloads. One identity, a landing zone, central logging and mandatory tags take a few weeks up front, and far longer if retrofitted.
Why cloud programmes disappoint
Cloud programmes rarely fail with an outage. They disappoint slowly: the bill beats the business case, the private platform is feared, or the migration stalls at 60 percent. The causes are familiar.
| What we find | What it leads to |
|---|---|
| Lift and shift with no right-sizing | VMs move as they were, sized for peak and running 24x7. The first annual bill is 1.5 to 2 times the estimate. |
| A renewal date driving the plan | The renewal quote arrives months before the date. The exit is rushed, or three more years are signed with no plan. |
| Accounts opened team by team | Nobody can say who has access to what, where the audit logs are, or whose costs are whose. |
| A private cloud nobody can run | Installed by a vendor and handed over without runbooks. Upgrades are postponed until support ends. |
| Data movement not costed | Applications that send large volumes out, or talk constantly across the link, cost more than their VMs. |
| Hybrid without design | Connectivity, DNS and identity were assumed rather than designed. Applications break in confusing ways after migration. |
Placement, cost and lock-in, explained properly
What makes a private cloud a cloud
A virtualised data centre is not a private cloud. The difference is self-service: teams create their own VMs, networks and volumes through a portal or an API, within quotas, without raising a ticket. Projects are isolated, usage is metered, and everything can be built from code. If you only need to run VMs reliably and nobody will use an API, a simpler KVM platform such as Proxmox VE is often the better choice.
Workload shape decides cost
Public cloud charges for every hour a resource exists. A VM that runs a few hours a day is cheap there. A database running 24x7 at steady load pays the provider’s margin every hour for five years. Most enterprise estates are dominated by the second kind.
Data gravity and egress
Public clouds charge little to bring data in and charge per gigabyte to send it out. Large datasets also take days or weeks to copy. Keep data close to the applications that use it most, and cost every flow that crosses a cloud boundary.
Four kinds of lock-in
| Kind | Example | How to limit it |
|---|---|---|
| Licence | Hypervisor subscription bundles, per-core database licences | Open hypervisor (KVM); check terms before placing databases |
| Data | Hundreds of terabytes in one provider’s object storage | Keep a copy or an exit plan for critical data |
| Service | Applications built on a provider-only queue, database or AI service | Use managed services where the value is worth the dependence |
| Skills and tooling | Everything built by clicking in one console | Infrastructure as code; open monitoring and logging tools |
The options compared
Most organisations end up using two or three of these models for different parts of the estate.
| Model | What it means | Strong at | Watch out for |
|---|---|---|---|
| Private cloud | Self-service cloud on hardware you own or lease | Steady workloads, data that must stay put, predictable cost over 3 to 5 years | Needs a team that can run and upgrade it; capacity must be bought ahead of demand |
| Public cloud | AWS, Azure, OCI or similar, by the hour or on 1 to 3 year commitments | Spiky or uncertain load, fast start, managed databases, analytics and AI services | Cost grows quietly; egress and always-on VMs add up; governance must be built |
| Hybrid cloud | Private and public joined by private links, one identity and DNS | Keeping core data at home while using public cloud where it is stronger | Connectivity, identity and DNS must be designed, not assumed |
| Multi-cloud | More than one public cloud | A service only one provider offers well, or a customer requirement | Every extra cloud multiplies the guardrails, tools and skills to maintain |
Choosing the private platform
For the private side, the right platform depends on how much self-service you need and what your team can operate.
| Platform | Best for | Trade-off |
|---|---|---|
| OpenStack on KVM with Ceph | Multi-tenant, API-driven cloud for several teams; 100 to several thousand VMs | The most capable option, and the most to learn; needs a disciplined upgrade process |
| Proxmox VE | Running VMs reliably for one IT team, from a handful of hosts to a few dozen | Simple to run; limited multi-tenancy and self-service API |
| OpenShift Virtualization | Teams standardised on Kubernetes wanting VMs and containers together | Subscription cost; VM operations follow Kubernetes patterns |
| A small remaining vSphere island | Appliances and vendor applications only supported on VMware | Keeps a licence, so keep it small and review it every year |
Deciding what runs where
Decide placement per application, with a one-line reason agreed by its owner (“private: steady load, payment data must stay in India”). Five questions cover almost every case:
- What is the shape of the load? Steady and 24x7 points to private. Seasonal, short-lived or fast-growing points to public.
- How much data moves, and where to? If it sends many terabytes a month out, cost the egress first. Keep chatty applications next to the systems they talk to.
- What do regulators and contracts require? Some data must stay in India, some on infrastructure you control, some needs audit rights over the provider.
- Does it rely on managed services? An application built on a provider’s managed database or AI service belongs there, unless you will rebuild it.
- Who will run it? A private platform needs Linux, networking and automation skills. Without them, public cloud or a managed private cloud is safer.
A typical estate comes out mixed:
| Workload | Typical placement | Main reason |
|---|---|---|
| ERP, core banking, hospital information system | Private | Steady 24x7 load, regulated data, heavy integration |
| Customer website and mobile back end | Public, or hybrid with data private | Spiky traffic, content delivery networks, fast release cycles |
| Data warehouse and analytics | Depends on data volume | Managed analytics is attractive; moving large data out is not |
| Development and test | Public, scheduled off out of hours | Needed about 60 hours a week, not 168 |
| Backup and long-term archive | Second site or public object storage | Cheap capacity, separate failure domain; cost the restore egress |
| GPU training bursts | Public or GPU cloud, then private for the steady base | Learn real usage before buying hardware |
Reference architecture
The diagram shows a hybrid estate during a VMware exit. A private cloud on OpenStack, KVM and Ceph runs the steady core; a public cloud landing zone runs what belongs there; both use one directory for sign-in. VMware shrinks until the last wave is signed off.
Numbered flows: (1) teams request VMs, networks and volumes through the portal, API or Terraform, signing in with their directory account; (2) the control plane places each VM on a compute node with room; (3) VM disks live on Ceph, not on the compute node, which makes live migration and fast recovery from a failed host possible; (4) backups go nightly to a separate, immutable target; (5) a private link such as Direct Connect, ExpressRoute or FastConnect joins the data centre to the cloud hub, with an IPsec VPN as backup; (6) VMs are converted from vSphere with virt-v2v and moved in waves.
Starting sizes for a first private cloud
| Role | Starting point | Why |
|---|---|---|
| Control nodes | 3 | The database and queue need a majority to survive one node failing or being upgraded. |
| Compute nodes | Sized from inventory, plus 1 spare | Lets you patch or lose one host without running out of room. |
| Ceph storage nodes | 3 minimum, 5 or more preferred | Three copies need three nodes; more nodes mean gentler recovery after a failure. |
| Networks | Management, storage, tenant overlay, external | Storage on its own network; 25 GbE to each server is a common baseline. |
Decisions that make or break it
- Choose the deployment tool for the next five years. Kolla-Ansible, OpenStack-Ansible and commercial distributions all work, if your team can upgrade with them.
- Design networks before installing anything. VLAN ranges, overlay networks, external IP ranges, MTU and an IP plan that does not clash with the cloud are painful to change later.
- Size Ceph for recovery, not just capacity. When a node fails, Ceph rebuilds its copies onto the remaining nodes. Leave free space and network capacity for that to happen without hurting users.
- Add the services teams expect. Load balancing (Octavia), DNS (Designate), secrets (Barbican) and instance high availability (Masakari).
- Decide who owns it. A private cloud is a product, not a project. Someone must own upgrades, capacity and on-call from day one.
Sizing and five-year cost
Compute
Size from the VM inventory, by processor and by memory, and take whichever needs more servers. Memory is rarely overcommitted; processor cores commonly are, at about three virtual CPUs per physical core.
Storage
Ceph keeps three copies of each block, and must keep enough free space to rebuild a failed node’s data onto the others before it reaches its 85 percent “near full” warning.
A worked example: 300 VMs
An estate of 300 VMs averages 4 vCPU, 16 GB of memory and 150 GB of disk each: 1,200 vCPU, 4,800 GB of memory and 45 TB of provisioned disk. Compute nodes have 64 physical cores and 768 GB of memory, of which 704 GB is usable after the host’s own needs.
| Item | Arithmetic | Result |
|---|---|---|
| Nodes by processor | 1,200 ÷ (64 × 3 × 0.75 = 144) | 8.3, so 9 nodes |
| Nodes by memory | 4,800 ÷ (704 × 0.85 = 598) | 8.0, so 9 nodes |
| Compute at go-live | 9 + 1 spare | 10 nodes, 12 by year three at 20% growth |
| Storage needed | 45 TB provisioned + 30% growth | About 58 TB usable |
| Ceph fill target, 6 nodes | 0.85 × 5 ÷ 6 | About 0.71 |
| Ceph raw capacity | 6 nodes × 12 × 3.84 TB NVMe = 276 TB; 276 ÷ 3 × 0.71 | About 65 TB usable |
Five-year cost, both ways
The same 300 VMs over five years. Planning ranges at typical Indian prices in 2026, not quotations.
| Cost line, five years | Private cloud (OpenStack, colocation) | Public cloud (3-year commitments) |
|---|---|---|
| Hardware: 3 control, 10 compute and 6 storage servers, network; 2 more compute nodes in year 3 | ₹3.5 to 5.5 crore | Included in hourly rates |
| Compute and storage consumption | Included above | ₹13 to 18 crore |
| Space and power: about 4 racks | ₹2.5 to 3.5 crore | Included |
| Software support, hardware maintenance | ₹1.5 to 3 crore | ₹0.7 to 1.5 crore (support plan) |
| Egress, NAT, load balancers, logging | Small | ₹1.5 to 3 crore |
| Backup, DR and network links | ₹1 to 2 crore | ₹1 to 2 crore |
| People: platform or cloud operations team | ₹4 to 5 crore (3 engineers) | ₹2.5 to 3.5 crore (2 engineers plus FinOps) |
| Five-year total | ₹12.5 to 19 crore | ₹18.7 to 28 crore |
On these assumptions the private cloud costs roughly a third less over five years. The gap narrows quickly if the private platform runs half empty, if the team must be hired at senior rates, or if right-sizing removes 20 to 30 percent of the public cloud capacity. Run the model with your own inventory.
Egress and licences: the lines most often missed
- Egress: at list prices of roughly ₹7 to 9 per GB, an application sending 20 TB a month out costs about 20,000 × ₹8 = ₹1.6 lakh a month, close to ₹20 lakh a year, before any VM is counted.
- Windows Server and SQL Server: per-core licences follow the physical cores of the host on any hypervisor. Check licence mobility rules before moving to public cloud.
- Oracle Database: licensing depends on how the hypervisor partitions processors; on many KVM platforms every host core can count. Check before a database moves.
- Hypervisor and cloud software: KVM, OpenStack and Ceph have no licence fee. A support subscription is optional but usually wise for production.
Landing zones, security and governance
A landing zone is the set of accounts, guardrails, networks and logging in place before the first workload arrives. Built first, from code, it takes a few weeks. Retrofitted after two years of ad hoc accounts, it is a project of its own. On the private cloud, projects and quotas play the role of accounts.
| Area | What good looks like |
|---|---|
| Identity | One sign-in through your directory, with roles in each cloud and in OpenStack. No local users or shared passwords; multi-factor everywhere; sealed break-glass accounts. |
| Account structure | Production, non-production and sandbox in separate accounts (projects on OpenStack), so a sandbox mistake cannot reach production. |
| Guardrails | Policies teams cannot switch off: approved regions, encryption and logging on, no public buckets, mandatory tags. |
| Network | Hub and spoke with a firewall and shared egress in the hub, and an IP plan that does not overlap with the data centre. |
| Logging | Audit logs and alerts from every account and cloud flow to one place teams cannot change. |
| Cost | Owner, environment and cost-centre tags enforced at creation; budgets and alerts per account. |
Regulation in practice
Indian rules shape placement more than architecture. RBI requires payment system data to be stored in India and expects audit and access rights over outsourced IT, including cloud; IRDAI and SEBI set similar expectations. CERT-In requires incidents to be reported within six hours and logs kept for 180 days within India. The DPDP Act sets duties on how personal data is processed and secured. Residency is achievable in Indian public cloud regions or a private cloud; the real questions are audit rights, log retention and who holds the keys.
Running it: day 2, upgrades and FinOps
Private clouds that disappoint were usually built well and then left without upgrades or capacity planning. Public cloud estates that disappoint never had an owner for cost. Both are fixed by a routine.
| Activity | How often | What good looks like |
|---|---|---|
| Health and alert review | Daily | Alerts that matter reach a person. Noise is tuned out, not ignored. |
| Security patching | Monthly | Rolling host patching with live migration; no user downtime. |
| Capacity review | Monthly | Forecast of when compute, storage and IP ranges run out, with time to buy. |
| Restore test | Monthly | One real restore, timed and recorded. |
| Cost review (public cloud) | Monthly | Idle and oversized resources removed; commitments matched to steady use. |
| Platform upgrade | Once a year | Rehearsed on staging, rolled out with a written rollback plan. |
Upgrades without fear
OpenStack releases every six months. Since the 2023.1 release, every other release supports an upgrade that skips one, so a yearly cycle is practical. Keep a small staging cloud built from the same code, rehearse each upgrade on it, and never fall more than one upgrade step behind. Platforms that skip upgrades for two or three years end up needing a migration instead.
FinOps for the public side
- Right-size first. Lifted VMs are often 30 to 50 percent larger than needed. Resize from four weeks of usage data before buying commitments.
- Commit to the steady base only. Reserved instances and savings plans cut 30 to 60 percent off compute for one or three years. Never commit to the peak.
- Schedule non-production. Off from 8 pm to 8 am and at weekends, development and test run 60 hours a week instead of 168, saving about 60 percent.
- Show cost by team every month. A report by owner tag changes behaviour faster than any policy.
- Watch data transfer. Egress, NAT gateway and cross-zone traffic are often the fastest-growing lines.
Roadmap: leaving VMware in planned waves
For a few hundred VMs, a private cloud build and VMware exit typically takes six to nine months. The pace is set by application owners’ testing time, not conversion speed.
| Stage | Typical duration | Outcome |
|---|---|---|
| 1. Assess | 3 to 4 weeks | Inventory (an RVTools export), dependency map, placement, five-year cost model, licence review. |
| 2. Design | 3 to 4 weeks | Architecture, network and IP plan, identity, bill of materials, wave plan, rollback rules. |
| 3. Build and prove | 6 to 10 weeks, after hardware arrives | Platform built from code and failure-tested; monitoring and backup working first. |
| 4. Pilot wave | 2 weeks | 10 to 20 low-risk VMs covering every operating system version. |
| 5. Production waves | 8 to 16 weeks | Waves of 20 to 60 VMs every one or two weeks, each tested side by side before cutover. |
| 6. Decommission | 2 to 4 weeks | Last sign-offs, VMware hosts retired, licences dropped. |
How a wave works
- Group by dependency. An application’s web, application and database VMs move together, so a cutover never splits a live dependency across a WAN link.
- Convert and test side by side. virt-v2v copies each VM, replaces VMware drivers with virtio drivers and removes VMware Tools. Large VMs are pre-copied and synced in the window to keep cutover short.
- Cut over only after sign-off. The application owner tests on KVM while the VMware copy is still live. Users move only when the owner agrees.
- Keep the way back. Source VMs stay intact and powered off for an agreed period, typically two to four weeks. Rollback is a DNS or IP change and a power-on.
What usually catches teams out
- Drivers and boot. Windows VMs need virtio drivers before or during conversion, or they will not boot. Test every operating system version in the pilot.
- Network identity. MAC addresses usually change. Anything licensed against a MAC address, and static DHCP reservations, need attention.
- Vendor support. Some appliances and commercial applications are supported only on VMware. Check before planning their wave; a few may stay on a small vSphere island or move to public cloud.
- Features and storage. Automatic load balancing like DRS is more limited on KVM, and Ceph behaves differently from VMFS and vSAN. Keep headroom and design volumes for the new platform.
Ten common mistakes
- A blanket cloud policy. “Cloud first” or “all private” puts some workloads in the wrong place by design.
- Comparing first-year cost. Five years, with people, egress, licences and refresh, is the only fair comparison.
- Lifting VMs without right-sizing. You pay every hour for capacity nobody uses.
- Letting the renewal date set the plan. Start the exit assessment at least 12 months before the VMware renewal.
- Choosing OpenStack when Proxmox VE would do, or the reverse. Pick the smallest platform that meets the need.
- Skipping upgrades. A private cloud three releases behind needs a migration, not an upgrade.
- Converting VMs alphabetically. Waves must follow dependencies, or a cutover splits a live application.
- Decommissioning VMware too early. Keep the source VMs until the owner signs off, and the hosts until the last wave is done.
- Two identities. Two audit trails and a clean-up project nobody planned.
- No owner for cost. Without tags, budgets and a monthly report by team, public cloud cost only goes one way.
Cloud readiness checklist (36 points)
Use this list to score your current position. Anything you cannot tick with evidence is a gap worth closing.
Strategy and placement
- Every application has a written placement and a one-line reason
- Placement agreed with each application owner
- Residency and audit requirements listed per application
- Dependencies mapped, so related systems are placed and moved together
- A named owner for the private platform and for public cloud cost
- Multi-cloud used only for a stated, specific need
The remaining 30 points are in the PDF, laid out as a printable checklist.
Get the full checklistGlossary
| Term | Meaning |
|---|---|
| Private cloud | Self-service, multi-tenant infrastructure on hardware you own or lease, driven by a portal and API. |
| OpenStack | Open-source cloud software: Nova (compute), Neutron (network), Cinder (block storage), Keystone (identity). |
| KVM | The hypervisor built into Linux, used by OpenStack, Proxmox VE and most public clouds. |
| Ceph | Open-source distributed storage providing block (RBD), object (RGW) and file storage across many servers. |
| OVN | Open Virtual Network: the software-defined networking used by Neutron. |
| virt-v2v | An open-source tool that converts VMware VMs to run on KVM, swapping in virtio drivers. |
| Landing zone | The accounts, guardrails, network and logging set up before workloads arrive in a cloud. |
| Egress | Data sent out of a public cloud, charged per gigabyte. |
| FinOps | The practice of making cloud cost visible to the teams who create it, and managing it monthly. |
About Vakratron Systems
Vakratron Systems is a vendor-neutral infrastructure design firm. We design data centre, disaster recovery, cloud, GPU and AI platforms for enterprises and government buyers, write our assumptions down, and stay with a design until it is running and tested.