Private cloud on OpenStack
OpenStack gives you AWS-style self-service on hardware you control: teams create their own VMs, networks and volumes through a portal or API, and you keep the data and the cost in your hands. Here is how a production platform is put together, and what decides whether it runs well.
How the pieces fit together
| Part | What it does |
|---|---|
| Control plane | Three nodes run the APIs, schedulers, database and message queue. Three is the minimum for the database and queue to keep a majority when one node is down. |
| 1 Scheduling | When a user asks for a VM, Nova picks a compute node with room, Neutron plugs it into the right network and Cinder attaches its disks. |
| 2 Storage traffic | VM disks live on Ceph, not on the compute node. That is what makes live migration and quick recovery from a failed host possible. |
| 3 Backup | Ceph keeps three copies, but replication is not backup. A separate backup target holds point-in-time copies. |
| 4 Outside access | OVN gateways route traffic between tenant networks and the internet, MPLS or the data centre core. |
Sizing a first platform
| Role | Starting point | Why |
|---|---|---|
| Control nodes | 3 | Database and queue need a majority to survive one failure. Can share nodes with networking at small scale. |
| Compute nodes | 3 or more, plus one spare | Lets you patch or lose one host without running out of capacity. |
| Ceph storage nodes | 3 minimum, 5 or more preferred | Three copies need three nodes. More nodes means faster recovery when a disk or node fails. |
| Networks | Management, storage, tenant overlay, external | Keeping storage traffic on its own network stops it from slowing down everything else. 25 GbE is a common baseline. |
Starting points only. Real sizing comes from your VM inventory, growth plans and how much failure you need to absorb.
Decisions that make or break it
- Choose the deployment tool for the next five years, not the next five weeks. Kolla-Ansible, OpenStack-Ansible and vendor distributions all work. What matters is that your team can upgrade with it.
- Plan upgrades from day one. OpenStack releases twice a year. Platforms that fall several releases behind become hard to upgrade. Schedule upgrades like patching, not like projects.
- Design networks before installing anything. VLAN ranges, overlay networks, external IP ranges and MTU are painful to change later.
- Size Ceph for recovery, not just capacity. When a node fails, Ceph rebuilds copies onto the remaining nodes. Leave enough free space and network capacity for that to happen without hurting users.
- Decide who runs it. A private cloud is a product, not a project. Someone must own upgrades, capacity and on-call.
Typical tools
| Layer | Common choices |
|---|---|
| Deployment | Kolla-Ansible, OpenStack-Ansible, Canonical OpenStack, Red Hat OpenStack Services on OpenShift |
| Hypervisor | KVM with libvirt |
| Networking | Neutron with OVN |
| Storage | Ceph (RBD for block, RGW for object) |
| Monitoring | Prometheus, Grafana, Alertmanager, OpenSearch or Loki for logs |
| Automation | Ansible, Terraform with the OpenStack provider |
Smaller environments that do not need multi-tenancy or a full API often do well with Proxmox VE instead. We will tell you if OpenStack is more than you need.
Need the same platform protected against a site failure? See disaster recovery on OpenStack.
Planning a cloud move or a VMware exit?
Send us a VM inventory export (RVTools is fine) and a line on what is driving the change. We will come back with a plain first view: what could move where, in what order, and roughly what it would cost to run.