A retailer that keeps its core at home and rents the peak
A retailer with about 350 stores and a growing online business had a simple problem with expensive consequences. On six or seven sale days a year the website needs ten times its normal capacity, and on the other 358 days it does not. This design keeps the systems that run the stores where they are and moves only the parts that need to stretch, with one identity, one network and stock that agrees on both sides.
A website that falls over on the days that matter most
The retailer sells apparel, footwear and home goods through about 350 stores and its own website and app. ERP, the POS back-end, inventory and order management run in the company’s own data centre, which was refreshed two years ago and still has plenty of life in it. The storefront runs there too, on a fixed set of virtual machines.
On a normal day the site handles around 8,000 concurrent shoppers. On the big sale days, festive season and end-of-season clearance, that rises to 80,000 or more in the first hours. Last year the site slowed to a crawl for most of the opening morning of the largest sale, and the promotions engine, which works out every discount at checkout, was the first thing to buckle.
Buying hardware for the peak would leave it idle for more than 350 days a year. Moving everything to public cloud would mean re-platforming an ERP that works well and feeds every till in the country. The question was how to give the website room to grow tenfold for a weekend, without overselling stock or creating a second set of everything to manage.
What could not be compromised
- ERP, POS back-end and inventory stay on-premises. Store operations cannot depend on an internet path.
- No overselling: stock shown online must be within a minute of what the stores and warehouses actually hold.
- Card data stays out of the retailer’s systems through a tokenised payment gateway, keeping PCI DSS scope small.
- Customer data is handled under the DPDP Act, with consent records and a clear view of where it is stored.
- Cloud spend must follow the sale calendar. Peak capacity is paid for in the sale weeks, not all year.
Four ways to handle a tenfold peak
Each option was costed over three years against last year’s traffic and the sale calendar for the next two years.
The storefront stretches, the core stays put
The storefront and promotions engine run on Kubernetes in a public cloud region in India, behind the CDN. They never call the ERP during a shopping session. Instead, they read from a local copy of the catalogue and stock, kept current by a stream of changes over a private link. Orders are queued in the cloud and drained into order management on-premises at a steady rate.
| Building block | Why it is there |
|---|---|
| 1 CDN and WAF | Serves images, static pages and cached product listings from the edge, and filters bots that hit stock pages on sale mornings. About 85% of requests never reach the storefront at all. |
| 2 Storefront on Kubernetes | The web front end and app APIs, packaged as containers. Runs at 6 pods on a normal day and scales to 60 on sale days, across three availability zones in an Indian region. |
| 3 Promotions engine | Works out every discount, coupon and bundle at basket and checkout. It runs next to the storefront so it can scale with it, and reads price rules that ERP publishes. |
| 4 Catalogue and stock cache | A read copy of products, prices and stock by location. Updated from change events, so the storefront never queries ERP or inventory directly during a session. |
| 5 Order queue and order management | Checkout writes each order to a durable queue in the cloud. Order management on-premises drains it at a rate it can handle, reserves stock and decides whether to ship from a warehouse or a store. |
| 6 Shared identity | One directory for staff and administrators across both environments, with SSO and MFA. Customer accounts sit in a customer identity service that uses the same policies and audit trail. |
| 7 Private link and one network design | Two 10G private connections to the cloud region on separate paths, one IP address plan, one firewall policy model and one DNS. The cloud is treated as another site, not as the internet. |
| 8 Cost controls | Every cloud resource is tagged by service and sale event. Budgets alert at 50%, 80% and 100%, and scale-down rules return capacity to baseline within hours of a sale closing. |
Sized from last year’s sale day, not from the average
Sizing used per-minute traffic from last year’s biggest sale, plus 30% growth. Normal days and sale days are sized separately because only the sale day drives the burst.
| Item | Normal day | Sale day peak | Basis |
|---|---|---|---|
| Concurrent shoppers | ~8,000 | ~80,000 | 10x measured, first two hours of the sale |
| Requests reaching storefront | ~225 per second | ~2,250 per second | ~1,500 and ~15,000 per second at the CDN, 85% served from cache |
| Storefront pods | 6 (2 per zone) | up to 60 | ~60 requests per second per pod, 2,250 / 60 = 38, plus 50% headroom |
| Orders | ~400 an hour | ~4,000 an hour | Queue drains at 6,000 an hour, so it never grows during the sale |
| Stock change events | ~10 lakh a day | ~30 lakh a day | ~2.5 lakh till transactions a day, about 4 lines each, 3x on sale days |
| Private link | ~80 Mbps used | ~300 Mbps used | 2 x 10G on separate paths, sized for resilience rather than volume |
| Cloud spend | ₹14 to 18 lakh a month | +₹5 to 8 lakh per sale event | Baseline storefront plus burst hours, at reserved and on-demand rates |
A full load test at 1.5 times the expected sale peak runs three weeks before each major sale, with the CDN, storefront, promotions, queue and order management all in the path.
From one shared network to the first live sale
The order of work follows the retail calendar. The first sale on the new design is a smaller one, and nothing changes in the four weeks before festive season.
Foundations
Weeks 1 to 6
Private links, IP plan, firewall policy, DNS and shared identity built and tested. Cloud landing zone with tagging and budgets from day one.
Gate: On-premises and cloud reach each other only over the private link, with every rule in one policy repository.
Data sync
Weeks 7 to 12
Change data capture on inventory and ERP price tables, streamed to the cloud cache. Stock accuracy measured against store counts every hour.
Gate: Cloud stock within 60 seconds of source for 99.9% of changes over two weeks.
Storefront and promotions
Weeks 13 to 18
Containers deployed to the cloud cluster, order queue built, order management connected. Traffic shifted by CDN weights: 5%, 25%, 100%.
Gate: Two weeks at 100% with checkout errors and page times equal or better than on-premises.
First sale
Weeks 19 to 22
Load test at 1.5x, autoscaling limits tuned, war room runbook rehearsed. A mid-size sale runs on the new design.
Gate: Sale completes with no slowdown, no oversold lines and spend inside the event budget.
Hand-over
Weeks 23 to 26
Runbooks, cost reports per sale and on-call rotas handed to the platform team. Old storefront VMs kept warm for one more sale, then retired.
Gate: Team runs a sale rehearsal on its own.
The risks in splitting a retailer across two places
| Risk | What could happen | How the design handles it |
|---|---|---|
| Overselling | Stock in the cloud lags a store sale and the last unit sells twice | Stock streams as changes within seconds, low-stock items keep a safety buffer online, and order management confirms the reservation before the confirmation email goes out. |
| Private link failure | The cloud loses its path to stock and order systems mid-sale | Two links on separate paths and providers. If both fail, the storefront keeps selling from the cache with wider safety buffers and orders wait in the queue. |
| Runaway spend | Autoscaling or a bot attack runs up a large bill | Hard upper limits on pod and node counts, budget alerts per event, and bot filtering at the CDN before traffic reaches anything that scales. |
| Promotion errors | A wrong price rule gives away margin at scale | Price rules come only from ERP, are checked against a margin floor before publishing, and can be withdrawn across the site in under a minute. |
| Two environments to run | The team ends up maintaining two of everything | One identity, one monitoring stack, one network policy repository and the same container pipeline for both sides. |
What was optimised
Served from the edge
Images, listings and static pages come from the CDN, so the storefront only handles what must be dynamic.
Capacity for the sale, not the year
Storefront and promotions scale from 6 to 60 pods and back within hours of the sale ending.
ERP calls per page
Shopping sessions read only the local cache. ERP and inventory never see sale-day traffic directly.
Stock freshness
Change data capture streams stock movements instead of overnight batches.
Network and identity design
The cloud region is another site on the same IP plan, firewall model and directory.
Cost visibility
Every sale has its own tags and budget, so the business sees what each event cost to run.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Sale-day site availability | Slowdown on the opening morning | Full service at 10x normal load, tested at 15x |
| Checkout page time (p95) | ~4 s on sale mornings | Under 1.5 s at sale peak |
| Oversold order lines | Several hundred per major sale | Close to zero, with safety buffers on low-stock items |
| Peak capacity cost | Hardware sized for peak, idle most of the year | Paid only in sale weeks, typically ₹5 to 8 lakh per event |
| Store operations | Depend on the data centre | Unchanged; no dependency on the cloud |
| Cost reporting | One IT budget line | Cost per sale event and per service |
Targets are validated in the load tests before each sale. Cloud spend depends on traffic, reserved capacity and pricing, and is re-modelled with your own sale calendar and traffic data during assessment.
What a team needs to deliver this
Hybrid network design
Private connectivity, IP planning and one firewall and DNS model across on-premises and cloud.
Kubernetes and autoscaling
Cluster design across zones, autoscaling limits and load testing for peak events.
Data sync and streaming
Change data capture, Kafka and cache design that keeps stock and prices consistent.
Identity
One directory, SSO and MFA for staff across environments, with customer identity alongside.
FinOps
Tagging, budgets and per-event cost reporting that the business can read.
Retail systems integration
ERP, POS, inventory and order management flows, including ship-from-store.
Other scenarios
Public cloud to private cloud
A B2B SaaS company moves its steady workloads off public cloud to a private cloud, and keeps burst capacity where it is cheap.
Read →Data centre consolidationThree data centres into one
A manufacturing group folds three ageing data centres from past acquisitions into one modern primary site and a DR site.
Read →Database consolidationOracle estate consolidation
An insurer consolidates 140 Oracle databases from 60 servers onto a private database platform, with fewer licensed cores and some databases moved to PostgreSQL.
Read →Does your website struggle on exactly the days you need it most?
Send us traffic from your last big sale and a rough map of what runs where. We will come back with a plain view of what should burst, what should stay put, and what a sale day would cost to run.