Two hundred small sites that keep working when the network does not
A building-materials manufacturer has 30 plants and 170 depots, many on the edge of towns where the network drops several times a week. Plant systems that wait on a central data centre stop when the link stops. This design puts a small Kubernetes cluster at every site so work carries on locally, and runs all 200 of them from one central private cloud, the same way, without anyone driving out to a depot to install an update.
When the link drops, the trucks stop
The company makes cement, ready-mix and building products at 30 plants and sells through 170 depots across several states. Weighbridges, dispatch, quality testing and MES interfaces at each site talk to applications in the head-office data centre. Plants have leased lines, depots mostly broadband, and both lose connectivity more often than anyone would like.
When a link goes down, dispatch stops: trucks queue at the weighbridge, challans are written by hand and keyed in later, sometimes twice. Plant quality results sit on a lab PC until someone copies them. Over the past year, network outages cost an estimated 1,100 site-hours of disrupted dispatch.
Some plants already run local servers, each set up differently by whoever was there at the time. Updates are done by remote desktop or by visit, and nobody can say with confidence which version of anything is running where. The CIO wanted sites that work on their own, but are managed as if they were one system.
What could not be compromised
- Every plant and depot must keep dispatching, weighing and recording quality results with the WAN down for at least 72 hours.
- No IT staff at depots and only a few at larger plants. Nothing may require a site visit for routine changes.
- Plant control networks stay separate from IT, following IEC 62443 zone principles. Edge clusters talk to PLCs only through gateways.
- Depot links are often 20 Mbps or less. Updates and data sync must work within that.
- Every site must run a known, recorded version of every application, provable at any time.
Four ways to run 200 small sites
The options were scored on what happens when the link goes down, how a change reaches 200 sites, and five-year cost.
Sites run themselves, the centre decides what they run
Each plant runs a three-node cluster and each depot a single-node cluster, with the same lightweight Kubernetes distribution and the same GitOps agent. The agent pulls its desired state from the central private cloud whenever the link is up, and keeps running the last known good state when it is not. Data is written locally first and forwarded when the network allows.
| Building block | Why it is there |
|---|---|
| 1 Edge Kubernetes at plants | Three small servers running a lightweight Kubernetes distribution, so one can fail without stopping the plant. Hosts MES connectors, dispatch, weighbridge integration, quality and a local historian. |
| 2 Single-node edge at depots | One server with the same stack and the same agent. If it fails, a pre-built spare is shipped and joins the fleet on power-up, pulling everything it needs from Git and the registry. |
| 3 Store-and-forward | Every transaction is written to a local queue first. When the link is up, events forward to the centre with a sequence number, so nothing is lost or counted twice. |
| 4 Fleet manager and Git | Git holds the desired state for every site: which applications, which versions, which settings. The fleet manager groups sites into rings and tracks which site runs what. |
| 5 Image registry | Signed and scanned container images. Plants keep a local registry cache, so a release is downloaded once per plant, and depots pull only changed layers. |
| 6 Central monitoring | Metrics and logs from all 260 edge nodes in one place, with alerts by site and by application. Sites buffer telemetry during outages, so gaps fill in afterwards. |
| 7 SD-WAN overlay | Two paths per site, leased line or broadband plus 4G, with central policy. Plant control networks are behind their own firewall and reach the edge cluster only through gateways. |
| 8 Identity and PKI | Each site has its own certificate and identity, issued at install. A stolen depot server can be cut off from the fleet in one step. |
Small at every site, sized for a week on its own
Site sizing comes from what each plant and depot runs today plus a margin. Central sizing comes from the number of nodes and the volume of data and telemetry they send back.
| Item | Figure | Basis |
|---|---|---|
| Edge clusters | 200 (30 plants + 170 depots) | One cluster per site |
| Edge nodes | 260 | 30 x 3 at plants + 170 x 1 at depots |
| Plant node | 16 cores, 64 GB, 2 x 1.92 TB SSD | 25 to 40 pods per plant, one node can fail |
| Local buffer | Plants 200 GB, depots 50 GB | ~2 GB a day per plant, ~150 MB per depot, 7 days plus generous margin |
| Release download | ~200 MB a typical change, ~1.5 GB full | ~80 seconds at 20 Mbps per depot, plants fetch once into their cache |
| Telemetry | ~5 lakh active series, ~200 GB logs a day | 260 nodes x ~2,000 series; 30 x 5 GB + 170 x 0.3 GB logs |
| Central platform | 12 nodes | 3 fleet and Git, 6 monitoring (30 days hot, ~6 TB), 3 Kafka brokers |
One plant and five depots run the full stack for a month before the hardware order for the rest of the fleet is placed, so the standard site sizes are based on measurements.
From one pilot plant to 200 sites, ring by ring
The rollout uses the same rings the fleet will use for every future change, so the team learns how to run the fleet while building it.
Design and central platform
Weeks 1 to 8
Central fleet manager, Git, registry, monitoring and ingest built. Site applications packaged as containers with their configuration in Git.
Gate: A test site is built from nothing using only Git and the registry.
Pilot
Weeks 9 to 14
One plant and five depots go live. The WAN is pulled deliberately for 72 hours at each, with dispatch running throughout.
Gate: Zero lost transactions after the outage tests, and data reconciled with the centre.
Ring 1
Weeks 15 to 20
Five plants and 30 depots in different regions, chosen to include the worst network links.
Gate: Fleet manager shows every site on the expected version; alerts reviewed weekly.
Rings 2 and 3
Weeks 21 to 38
The remaining 24 plants and 134 depots, about 3 plants and 15 depots a fortnight. Spares pool set up in each region.
Gate: Each fortnight’s sites stable for a week before the next batch ships.
Hand-over
Weeks 39 to 42
Runbooks, on-call and change process handed to the platform team. Old local servers retired.
Gate: Team ships a change to all 200 sites through the rings on its own.
What could go wrong with 200 sites
| Risk | What could happen | How the design handles it |
|---|---|---|
| A bad change reaches every site | A faulty release stops dispatch everywhere at once | Changes go to five canary sites first, then 25%, then the rest, with automatic halts on failed health checks. |
| Long outages | A site is offline for longer than the buffer allows | Buffers sized for 7 days against a 72-hour target, with alerts at 50% full. 4G backup on every site. |
| Duplicate or lost transactions | Forwarded events are replayed twice or dropped | Every event carries a site sequence number. Central ingest de-duplicates and reports gaps per site. |
| Physical theft or tampering | A depot server is stolen or tampered with | Disks encrypted, site certificates revocable centrally, and no secrets that unlock other sites. |
| Plant network exposure | Edge cluster becomes a path into plant control systems | Edge talks to PLCs only through gateways in a separate zone, with one-way flows where the process allows. |
What was optimised
Offline operation
Sites keep dispatching, weighing and recording quality results without the WAN.
One change, every site
A change to 200 sites is a reviewed commit, rolled out ring by ring.
Standard site designs
One plant design and one depot design, so spares and runbooks are shared.
Downloads per plant
Local registry caches mean a release crosses the plant link once.
Site visits for updates
Routine changes and even rebuilds need only a power cable and a network link.
Fleet status
Versions, health and buffer levels for every site on one screen.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Dispatch during WAN outages | Stops, paper challans | Continues normally for at least 72 hours |
| Site-hours of disrupted dispatch | ~1,100 a year | Down 90% or more |
| Time to roll a change to all sites | Weeks, with visits | 1 to 3 days through the rings |
| Sites on a known version | Unknown | All 200, visible at any time |
| Data re-keyed after outages | Common | None; forwarded automatically |
| Replacing a failed depot server | Days plus a visit | Next-day spare, rebuilt from Git on power-up |
Targets are confirmed in the pilot, including deliberate outage tests. Results depend on how many site applications can be packaged as containers, which is assessed application by application.
What a team needs to deliver this
Edge Kubernetes
Lightweight distributions, small-cluster design and behaviour when the network disappears.
GitOps at fleet scale
Pull-based delivery, rollout rings, automatic halts and drift detection across hundreds of clusters.
Data sync
Store-and-forward queues, sequence numbers and de-duplication that survive long outages.
SD-WAN design
Dual-path site connectivity, central policy and bandwidth planning for small links.
OT and IT separation
Zone design between plant control networks and edge platforms.
Observability
Fleet-wide metrics and logs, with alerts that make sense per site.
Other scenarios
Public cloud to private cloud
A B2B SaaS company moves its steady workloads off public cloud to a private cloud, and keeps burst capacity where it is cheap.
Read →Hybrid cloudHybrid cloud for omnichannel retail
An omnichannel retailer keeps ERP and stores on-premises and bursts its online storefront to public cloud for 10x sale-day traffic.
Read →Data centre consolidationThree data centres into one
A manufacturing group folds three ageing data centres from past acquisitions into one modern primary site and a DR site.
Read →Running many small sites that stop when the network does?
Tell us how many sites you have, what runs at each and how good the links are. We will come back with a plain view of what should run at the edge, what should stay central and how the fleet would be run.