Home / Kubernetes / Platform design
Kubernetes guide

Platform design

A good platform makes the right way the easy way: developers get a working, secure environment from a template, deploy by merging to Git, and see their own metrics. Here is how such a platform is put together, and the early decisions that matter most.

Control plane3 nodes
DeliveryGitOps
TenancyNamespace per team
Upgrades~3 Kubernetes releases a year
Reference architecture

How the pieces fit together

Scroll sideways to see the whole diagram →
Platform servicesWorkload clusterDeveloperspush code, use templatesDeveloper portaltemplates, docs, self-serviceCI pipelinesbuild, test, scanContainer registryscanned and signed imagesGitOps controllerArgo CD or FluxIngress and gatewayTLS, routing, rate limitsObservabilitymetrics, logs, tracesTeam namespacesquotas and network policies per teamCluster configurationpolicies, add-ons, secrets operator1234
Developers never touch the cluster directly. Code goes through CI, images through the registry, and configuration through Git.
PartWhat it does
1 BuildCI builds the container image, runs tests and scans it for known vulnerabilities.
2 ImagesOnly scanned, signed images from your registry are allowed to run.
3 Configuration from GitThe GitOps controller applies everything the cluster runs from Git, including policies and add-ons.
4 FeedbackEach team sees its own metrics, logs and alerts, so problems are found by the people who can fix them.
Team namespacesEach team gets a namespace with resource quotas and network policies, created from a template.

Managed or self-run?

Managed Kubernetes (EKS, AKS, GKE, OKE)Self-run (RKE2, OpenShift, kubeadm)
Control planeRun and upgraded by the providerRun and upgraded by you
Good forPublic cloud workloads, small platform teamsOn-premises, private cloud, strict data rules
Watch out forCloud cost and provider-specific featuresUpgrade discipline and on-call ownership

Sizing a first cluster

  • Control plane: three nodes, so it survives one failure and can be upgraded one node at a time.
  • Workers: enough capacity for peak load with one node lost, plus room for rolling updates.
  • Node pools: separate pools for general workloads, memory-heavy workloads and GPUs if needed.
  • Storage: a CSI driver for your storage (Ceph via Rook, vendor arrays, or cloud volumes) for stateful services.

Planning a Kubernetes platform, or fixing one?

Tell us how many applications and teams you have, how you deploy today and where it should run. We will come back with an honest view: whether Kubernetes fits, what a lean platform looks like, and what it takes to run.