Three ageing data centres, one modern home
A manufacturing group grew by buying companies, and with each one it inherited a server room. Ten years on, it runs three data centres with three hypervisors, seven storage arrays and three directories, and pays to cool all of them. This design brings everything into one primary site and one recovery site, wave by wave, without stopping a single production line.
Three of everything, and none of it young
The group makes industrial components at 24 plants. Its IT runs from three sites: a server room at head office on VMware, a plant-campus data centre on Hyper-V that came with the first acquisition, and a smaller room at the second acquired unit with a mix of an older KVM stack and about 35 physical servers. Together they host about 900 virtual machines and 120 physical servers.
Most of the hardware is between six and eleven years old. Two of the storage arrays are past vendor support, the plant-campus site has had two cooling failures in eighteen months, and the head-office room cannot take more power. Each site has its own firewall rules, its own backup tool and its own Active Directory domain, and nobody has a complete list of what depends on what.
The board approved a refresh, but refreshing three sites would mean buying three of everything again. The CIO’s question was whether the group could end up with one well-run primary site and a proper recovery site, for less than a like-for-like refresh, and move without disrupting plants that run around the clock.
What could not be compromised
- Plants run three shifts, seven days a week. ERP, MES interfaces and dispatch systems cannot be down for more than a planned four-hour window, once a month.
- Applications that talk to each other must move together, or their traffic must be proven to cope with the distance during the move.
- About 20 systems are tied to physical hardware: licence dongles, per-socket licences or vendor certification.
- The IT team is 14 people across three sites, used to different tools. The end state must be something one team can run.
- Year-end close and the annual plant maintenance shutdown are fixed. Migration waves fit around them.
Four ways to fix three ageing sites
Each option was costed over seven years, including hardware, facilities, power, people and migration effort.
One primary site, one recovery site, one way of running both
The primary site runs one virtualisation platform, one storage platform and one network fabric, with a small bare-metal pool for systems that must stay physical. The recovery site receives storage replicas and a second backup copy, and holds warm compute that takes over in a disaster. The three old sites drain into the new one in dependency waves.
| Building block | Why it is there |
|---|---|
| 1 WAN and SD-WAN hub | Every plant and office connects to the primary site over two paths, a leased line and broadband, with the recovery site as a second hub. Replaces three separate WAN designs. |
| 2 Edge firewalls | One active-standby pair with one rule set, rebuilt from the three old rule bases. About 60% of the old rules were unused or duplicated and are not carried over. |
| 3 Spine-leaf fabric | Two 100G spines and 25G to every host. Storage, VM and management traffic on separate networks, designed so a link or switch failure does not stop anything. |
| 4 Virtualisation cluster | Sixteen hosts on one hypervisor replace three platforms. VMs are converted during migration, so the team learns one tool and one set of runbooks. |
| 5 Bare-metal pool | Fourteen current servers for the roughly 20 systems that must stay physical, combined where licences allow. Each one documented with why it cannot be virtualised. |
| 6 All-flash storage | One array with deduplication and compression replaces seven. Snapshots feed backup and asynchronous replication feeds the recovery site. |
| 7 One directory | The three Active Directory domains are merged into one, with group policies cleaned up. Done early, because almost every migration depends on it. |
| 8 Recovery site | Storage replica, warm compute for the critical tier and a second, offline backup copy. Failover is tested twice a year with real applications. |
Sized from what the estate actually uses
Ninety days of performance data was collected from all three sites. Most VMs were sized years ago and never revisited, so allocated capacity is far above real use.
| Resource | Three sites today | New primary site | Basis |
|---|---|---|---|
| vCPU | ~3,600 allocated, ~1,300 in use at p95 | 16 hosts, 2 x 32 cores each, 1,024 cores | Right-sized to ~2,700 vCPU at 4:1 = ~675 cores, +30% growth, N+2 hosts |
| Memory | ~10.8 TB allocated, ~7 TB in use | 16 x 1 TB = 16 TB | ~8.5 TB right-sized, +30% growth, still fits with 2 hosts down |
| Physical servers | 120 | 14 bare-metal servers | ~70 virtualised, ~30 retired, ~20 stay physical on newer hardware |
| Storage | 7 arrays, ~420 TB raw, ~210 TB used | 1 all-flash array, ~110 TB usable | ~60 TB cold data to archive, 150 TB active +40% growth = 210 TB, at 2:1 reduction |
| Racks | ~46 across three rooms | 12 primary, 6 DR | Measured power per rack, 12 kW design |
| IT power | ~310 kW, PUE about 2.0 | ~150 kW primary + ~40 kW DR, PUE about 1.5 | ~620 kW at the meter today, ~285 kW after |
| WAN | 3 hub designs | 1 SD-WAN hub per site, 2 x 1G internet + leased line | Plant traffic measured at peak shift change |
The power figures use measured draw from the three sites, not nameplate ratings. Colocation contracts are written for 12 racks with an option for 4 more, so growth does not mean a new site.
Six waves, planned around how applications really talk to each other
Waves are built from observed network traffic between systems, not from the server list. Systems that talk to each other a lot move together. Nothing moves during year-end close or the plant shutdown.
Discover and plan
Weeks 1 to 8
Traffic and performance collected from all three sites. Dependency groups built, owners confirmed, licence-bound systems identified. Colocation contracts signed and hardware ordered.
Gate: Every system has an owner, a wave, a move method and a rollback plan.
Build the primary site
Weeks 9 to 18
Racks, fabric, firewalls, cluster, storage and backup built from code. Directory merge done. SD-WAN hub live with plants dual-homed to old and new hubs.
Gate: Failure tests passed: host, switch, storage controller and WAN link.
Pilot wave
Weeks 19 to 22
About 60 low-risk VMs, including file servers and internal tools, move from all three sites. Tooling and runbooks tuned.
Gate: No incidents for two weeks, measured move time within plan.
Production waves 2 to 6
Weeks 23 to 40
Dependency groups move in monthly windows, the largest risk last. ERP and MES interfaces move together in one planned four-hour window.
Gate: Each wave stable for 14 days before the next starts.
Build DR and close sites
Weeks 41 to 48
Recovery site live with replicas, first full failover test. Old hardware wiped with certificates, rooms handed back or repurposed.
Gate: Failover test passed, three sites empty, power contracts closed.
What could go wrong in a three-way move
| Risk | What could happen | How the design handles it |
|---|---|---|
| Hidden dependencies | A system moves and breaks one nobody knew it talked to | Waves built from 90 days of observed traffic, not from documentation. Anything unexplained gets an owner before its wave is scheduled. |
| Latency between sites mid-move | An application is split across old and new sites and slows down | Dependency groups move together. Where that is impossible, the inter-site link is load-tested with the real application first. |
| Plant disruption | MES or dispatch interfaces fail during a shift | Plant-facing systems move only in agreed four-hour windows, with plant IT on site and a rehearsed rollback. |
| Old hardware failing before its turn | An out-of-support array fails mid-programme | The two oldest arrays are in the first production waves, and their backups are copied to the new site in the build phase. |
| One team, three habits | Staff keep running things the old way | One runbook set, training from the build phase, and the team does the pilot wave themselves with support alongside. |
Where the savings come from
Lower power draw
From about 620 kW at the meter to about 285 kW across both new sites, through newer servers and better cooling.
Fewer racks
Right-sized hosts and one storage array fit in twelve racks at the primary site and six at DR.
One of everything
One hypervisor, one storage platform, one firewall policy, one directory and one backup tool.
Servers retired
Systems with no users or a modern replacement are switched off, not migrated.
Fewer firewall rules
Unused and duplicate rules from three old rule bases are not carried over.
Storage reduction
Deduplication and compression on all-flash, with cold data moved to cheaper archive.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Facility power at the meter | ~620 kW | ~285 kW, saving roughly ₹2.3 to 2.9 crore a year at ₹8 to ₹10 per kWh |
| Seven-year cost | Like-for-like refresh of three sites | 25 to 35% lower, including colocation fees |
| Recovery | Backups only, restore in days | RPO 15 minutes, RTO 4 hours for critical systems, tested twice a year |
| Platforms to run | 3 hypervisors, 7 arrays, 3 domains | 1 of each |
| Hardware in support | About half | All, with a five-year support term |
| Change to an application | Days of finding where it runs | One CMDB, one monitoring view |
Targets are confirmed after discovery. Power savings depend on measured draw, colocation pricing and the final number of systems retired, and are recalculated with your own meter data during assessment.
What a team needs to deliver this
Discovery and dependency mapping
Building migration waves from observed traffic and real usage, not from spreadsheets.
Data centre and colocation design
Rack, power, cooling and Tier III colocation selection and contracts.
Network design
Spine-leaf fabrics, SD-WAN hub design and firewall rule rationalisation.
Virtualisation migration
Moving VMs between hypervisors, P2V and conversion at scale with rollback.
Storage consolidation
Moving many arrays onto one platform, replication and tiering cold data.
DR design and testing
Recovery tiers, replication and failover tests that use real applications.
Other scenarios
Public cloud to private cloud
A B2B SaaS company moves its steady workloads off public cloud to a private cloud, and keeps burst capacity where it is cheap.
Read →Hybrid cloudHybrid cloud for omnichannel retail
An omnichannel retailer keeps ERP and stores on-premises and bursts its online storefront to public cloud for 10x sale-day traffic.
Read →Database consolidationOracle estate consolidation
An insurer consolidates 140 Oracle databases from 60 servers onto a private database platform, with fewer licensed cores and some databases moved to PostgreSQL.
Read →Running more data centres than your business needs?
Send us a rough inventory of your sites, hardware ages and power bills. We will come back with a plain first view of what consolidation could look like, what it would cost and where the savings would come from.