Home / Cloud / Private cloud on OpenStack
Cloud guide

Private cloud on OpenStack

OpenStack gives you AWS-style self-service on hardware you control: teams create their own VMs, networks and volumes through a portal or API, and you keep the data and the cost in your hands. Here is how a production platform is put together, and what decides whether it runs well.

HypervisorKVM
Smallest sensible start3 control + 3 compute + 3 storage nodes
Typical storageCeph, three copies
Best forSteady workloads, data that stays put
Reference architecture

How the pieces fit together

Scroll sideways to see the whole diagram →
Control plane (3 nodes)Compute, storage and networkUsers and teamsself-service portal and APIAPI and dashboardHorizon, Keystone, service APIsCloud servicesNova, Neutron, Cinder, GlanceDatabase and message queueMariaDB Galera, RabbitMQMonitoring and logsPrometheus, Grafana, OpenSearchCompute nodesKVM hypervisors, add more as you growCeph storage clusterblock, image and object storageBackup targetseparate from CephNetwork gatewaysOVN routing to outside networksOutside networksinternet, MPLS, data centre core1234
Numbered lines show the main flows. The control plane runs on three nodes so that one can fail or be upgraded without stopping the cloud.
PartWhat it does
Control planeThree nodes run the APIs, schedulers, database and message queue. Three is the minimum for the database and queue to keep a majority when one node is down.
1 SchedulingWhen a user asks for a VM, Nova picks a compute node with room, Neutron plugs it into the right network and Cinder attaches its disks.
2 Storage trafficVM disks live on Ceph, not on the compute node. That is what makes live migration and quick recovery from a failed host possible.
3 BackupCeph keeps three copies, but replication is not backup. A separate backup target holds point-in-time copies.
4 Outside accessOVN gateways route traffic between tenant networks and the internet, MPLS or the data centre core.

Sizing a first platform

RoleStarting pointWhy
Control nodes3Database and queue need a majority to survive one failure. Can share nodes with networking at small scale.
Compute nodes3 or more, plus one spareLets you patch or lose one host without running out of capacity.
Ceph storage nodes3 minimum, 5 or more preferredThree copies need three nodes. More nodes means faster recovery when a disk or node fails.
NetworksManagement, storage, tenant overlay, externalKeeping storage traffic on its own network stops it from slowing down everything else. 25 GbE is a common baseline.

Starting points only. Real sizing comes from your VM inventory, growth plans and how much failure you need to absorb.

Decisions that make or break it

  • Choose the deployment tool for the next five years, not the next five weeks. Kolla-Ansible, OpenStack-Ansible and vendor distributions all work. What matters is that your team can upgrade with it.
  • Plan upgrades from day one. OpenStack releases twice a year. Platforms that fall several releases behind become hard to upgrade. Schedule upgrades like patching, not like projects.
  • Design networks before installing anything. VLAN ranges, overlay networks, external IP ranges and MTU are painful to change later.
  • Size Ceph for recovery, not just capacity. When a node fails, Ceph rebuilds copies onto the remaining nodes. Leave enough free space and network capacity for that to happen without hurting users.
  • Decide who runs it. A private cloud is a product, not a project. Someone must own upgrades, capacity and on-call.

Typical tools

LayerCommon choices
DeploymentKolla-Ansible, OpenStack-Ansible, Canonical OpenStack, Red Hat OpenStack Services on OpenShift
HypervisorKVM with libvirt
NetworkingNeutron with OVN
StorageCeph (RBD for block, RGW for object)
MonitoringPrometheus, Grafana, Alertmanager, OpenSearch or Loki for logs
AutomationAnsible, Terraform with the OpenStack provider

Smaller environments that do not need multi-tenancy or a full API often do well with Proxmox VE instead. We will tell you if OpenStack is more than you need.

Need the same platform protected against a site failure? See disaster recovery on OpenStack.

Planning a cloud move or a VMware exit?

Send us a VM inventory export (RVTools is fine) and a line on what is driving the change. We will come back with a plain first view: what could move where, in what order, and roughly what it would cost to run.