Home / Deployment scenarios / Edge clusters run from a central cloud
Reference deployment Edge and Kubernetes

Two hundred small sites that keep working when the network does not

A building-materials manufacturer has 30 plants and 170 depots, many on the edge of towns where the network drops several times a week. Plant systems that wait on a central data centre stop when the link stops. This design puts a small Kubernetes cluster at every site so work carries on locally, and runs all 200 of them from one central private cloud, the same way, without anyone driving out to a depot to install an update.

SectorBuilding materials
Sites30 plants, 170 depots
Programme10 months, ring by ring
ModelEdge clusters + central GitOps
The situation

When the link drops, the trucks stop

The company makes cement, ready-mix and building products at 30 plants and sells through 170 depots across several states. Weighbridges, dispatch, quality testing and MES interfaces at each site talk to applications in the head-office data centre. Plants have leased lines, depots mostly broadband, and both lose connectivity more often than anyone would like.

When a link goes down, dispatch stops: trucks queue at the weighbridge, challans are written by hand and keyed in later, sometimes twice. Plant quality results sit on a lab PC until someone copies them. Over the past year, network outages cost an estimated 1,100 site-hours of disrupted dispatch.

Some plants already run local servers, each set up differently by whoever was there at the time. Updates are done by remote desktop or by visit, and nobody can say with confidence which version of anything is running where. The CIO wanted sites that work on their own, but are managed as if they were one system.

What could not be compromised

  • Every plant and depot must keep dispatching, weighing and recording quality results with the WAN down for at least 72 hours.
  • No IT staff at depots and only a few at larger plants. Nothing may require a site visit for routine changes.
  • Plant control networks stay separate from IT, following IEC 62443 zone principles. Edge clusters talk to PLCs only through gateways.
  • Depot links are often 20 Mbps or less. Updates and data sync must work within that.
  • Every site must run a known, recorded version of every application, provable at any time.
Options weighed

Four ways to run 200 small sites

The options were scored on what happens when the link goes down, how a change reaches 200 sites, and five-year cost.

OptionWhat worksWhat does notVerdict
Keep everything central, improve the WANSimple to run, one copy of each application.Even dual links fail at rural depots. Every outage still stops dispatch.Rejected
Local servers or VMs at each site, managed by handWorks offline. Familiar.Two hundred hand-built sites, drift between them, and visits for every change.Rejected
Public-cloud managed edge serviceFleet tooling included.Per-site fees across 200 sites, and management depends on an internet path to the provider.Rejected
Lightweight Kubernetes at every site, GitOps from a central private cloudSites run on their own. Every site is defined in Git and updated the same way, ring by ring.Needs disciplined packaging of site applications and a small central platform team.Chosen
Target architecture

Sites run themselves, the centre decides what they run

Each plant runs a three-node cluster and each depot a single-node cluster, with the same lightweight Kubernetes distribution and the same GitOps agent. The agent pulls its desired state from the central private cloud whenever the link is up, and keeps running the last known good state when it is not. Data is written locally first and forwarded when the network allows.

Scroll sideways to see the whole diagram →
CENTRAL PRIVATE CLOUDNETWORKPLANT EDGE, 30 SITESDEPOT EDGE, 170 SITESIdentity and PKIsite certificates, SSOGit repositorydesired state for 200 sitesImage registrysigned images, scannedCentral data platformproduction and dispatch dataRollout ringscanary, 25%, then allFleet managerGitOps for 200 clustersCentral monitoringmetrics, logs, alertsEvent ingestKafka, de-duplicationSD-WAN overlayleased line or broadband plus 4G backup at every site, encrypted, central policyEdge Kubernetes3 nodes, GitOps agentStore-and-forwardlocal queue, 7 daysPlant floorMES, PLC gateways, scalesLocal servicesDNS, registry and auth cacheSingle-node edgesame stack and agent, one serverDepot appsdispatch, stock, weighbridgeimages2pullpull3forwardforward154Control / API callScheduled copyLogging / managementData / replication
Numbered flows: (1) a change is committed to Git and picked up by the fleet manager, (2) site agents pull their desired state over SD-WAN, ring by ring, (3) site data is queued locally and forwarded when the link is up, (4) forwarded events are de-duplicated in central ingest and land in the data platform, (5) every site reports metrics and logs to central monitoring.
Building blockWhy it is there
1 Edge Kubernetes at plantsThree small servers running a lightweight Kubernetes distribution, so one can fail without stopping the plant. Hosts MES connectors, dispatch, weighbridge integration, quality and a local historian.
2 Single-node edge at depotsOne server with the same stack and the same agent. If it fails, a pre-built spare is shipped and joins the fleet on power-up, pulling everything it needs from Git and the registry.
3 Store-and-forwardEvery transaction is written to a local queue first. When the link is up, events forward to the centre with a sequence number, so nothing is lost or counted twice.
4 Fleet manager and GitGit holds the desired state for every site: which applications, which versions, which settings. The fleet manager groups sites into rings and tracks which site runs what.
5 Image registrySigned and scanned container images. Plants keep a local registry cache, so a release is downloaded once per plant, and depots pull only changed layers.
6 Central monitoringMetrics and logs from all 260 edge nodes in one place, with alerts by site and by application. Sites buffer telemetry during outages, so gaps fill in afterwards.
7 SD-WAN overlayTwo paths per site, leased line or broadband plus 4G, with central policy. Plant control networks are behind their own firewall and reach the edge cluster only through gateways.
8 Identity and PKIEach site has its own certificate and identity, issued at install. A stolen depot server can be cut off from the fleet in one step.
Sizing, worked out

Small at every site, sized for a week on its own

Site sizing comes from what each plant and depot runs today plus a margin. Central sizing comes from the number of nodes and the volume of data and telemetry they send back.

ItemFigureBasis
Edge clusters200 (30 plants + 170 depots)One cluster per site
Edge nodes26030 x 3 at plants + 170 x 1 at depots
Plant node16 cores, 64 GB, 2 x 1.92 TB SSD25 to 40 pods per plant, one node can fail
Local bufferPlants 200 GB, depots 50 GB~2 GB a day per plant, ~150 MB per depot, 7 days plus generous margin
Release download~200 MB a typical change, ~1.5 GB full~80 seconds at 20 Mbps per depot, plants fetch once into their cache
Telemetry~5 lakh active series, ~200 GB logs a day260 nodes x ~2,000 series; 30 x 5 GB + 170 x 0.3 GB logs
Central platform12 nodes3 fleet and Git, 6 monitoring (30 days hot, ~6 TB), 3 Kafka brokers

One plant and five depots run the full stack for a month before the hardware order for the rest of the fleet is placed, so the standard site sizes are based on measurements.

How it is delivered

From one pilot plant to 200 sites, ring by ring

The rollout uses the same rings the fleet will use for every future change, so the team learns how to run the fleet while building it.

1

Design and central platform

Weeks 1 to 8

Central fleet manager, Git, registry, monitoring and ingest built. Site applications packaged as containers with their configuration in Git.

Gate: A test site is built from nothing using only Git and the registry.

2

Pilot

Weeks 9 to 14

One plant and five depots go live. The WAN is pulled deliberately for 72 hours at each, with dispatch running throughout.

Gate: Zero lost transactions after the outage tests, and data reconciled with the centre.

3

Ring 1

Weeks 15 to 20

Five plants and 30 depots in different regions, chosen to include the worst network links.

Gate: Fleet manager shows every site on the expected version; alerts reviewed weekly.

4

Rings 2 and 3

Weeks 21 to 38

The remaining 24 plants and 134 depots, about 3 plants and 15 depots a fortnight. Spares pool set up in each region.

Gate: Each fortnight’s sites stable for a week before the next batch ships.

5

Hand-over

Weeks 39 to 42

Runbooks, on-call and change process handed to the platform team. Old local servers retired.

Gate: Team ships a change to all 200 sites through the rings on its own.

Way back: Every change is a Git commit, so rolling back is reverting the commit. Rings stop automatically if health checks fail at canary sites, and a site that cannot reach the centre keeps running its last known good version.
Risks, handled up front

What could go wrong with 200 sites

RiskWhat could happenHow the design handles it
A bad change reaches every siteA faulty release stops dispatch everywhere at onceChanges go to five canary sites first, then 25%, then the rest, with automatic halts on failed health checks.
Long outagesA site is offline for longer than the buffer allowsBuffers sized for 7 days against a 72-hour target, with alerts at 50% full. 4G backup on every site.
Duplicate or lost transactionsForwarded events are replayed twice or droppedEvery event carries a site sequence number. Central ingest de-duplicates and reports gaps per site.
Physical theft or tamperingA depot server is stolen or tampered withDisks encrypted, site certificates revocable centrally, and no secrets that unlock other sites.
Plant network exposureEdge cluster becomes a path into plant control systemsEdge talks to PLCs only through gateways in a separate zone, with one-way flows where the process allows.
What was optimised

What was optimised

72 h+

Offline operation

Sites keep dispatching, weighing and recording quality results without the WAN.

1 commit

One change, every site

A change to 200 sites is a reviewed commit, rolled out ring by ring.

2 shapes

Standard site designs

One plant design and one depot design, so spares and runbooks are shared.

Once

Downloads per plant

Local registry caches mean a release crosses the plant link once.

0

Site visits for updates

Routine changes and even rebuilds need only a power cable and a network link.

1 view

Fleet status

Versions, health and buffer levels for every site on one screen.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
Dispatch during WAN outagesStops, paper challansContinues normally for at least 72 hours
Site-hours of disrupted dispatch~1,100 a yearDown 90% or more
Time to roll a change to all sitesWeeks, with visits1 to 3 days through the rings
Sites on a known versionUnknownAll 200, visible at any time
Data re-keyed after outagesCommonNone; forwarded automatically
Replacing a failed depot serverDays plus a visitNext-day spare, rebuilt from Git on power-up

Targets are confirmed in the pilot, including deliberate outage tests. Results depend on how many site applications can be packaged as containers, which is assessed application by application.

Skills this draws on

What a team needs to deliver this

Edge Kubernetes

Lightweight distributions, small-cluster design and behaviour when the network disappears.

GitOps at fleet scale

Pull-based delivery, rollout rings, automatic halts and drift detection across hundreds of clusters.

Data sync

Store-and-forward queues, sequence numbers and de-duplication that survive long outages.

SD-WAN design

Dual-path site connectivity, central policy and bandwidth planning for small links.

OT and IT separation

Zone design between plant control networks and edge platforms.

Observability

Fleet-wide metrics and logs, with alerts that make sense per site.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

Running many small sites that stop when the network does?

Tell us how many sites you have, what runs at each and how good the links are. We will come back with a plain view of what should run at the edge, what should stay central and how the fleet would be run.