Home / Deployment scenarios / MLOps platform for demand forecasting
Reference deployment MLOps

From notebooks on laptops to forecasts that retrain themselves, with a person still signing off

An FMCG company had good data scientists and models that worked, once. Forecasts were rebuilt by hand each quarter on laptops, emailed as spreadsheets, and quietly went stale in between. This design puts the whole model lifecycle on one platform, so every forecast can be traced to its data and code, drift is caught in weeks rather than quarters, and nothing reaches planners without a named person approving it.

SectorFMCG, foods and home care
Scope2,000 SKUs x 25 regions
Programme24 weeks
ModelOpen-source MLOps on Kubernetes
The situation

Models that worked in the notebook and faded in production

The company sells about 2,000 SKUs through 25 distributor regions, with primary sales in the ERP and secondary sales coming in daily from distributors’ DMS feeds. A data science team of six builds a weekly demand forecast for every SKU and region, about 50,000 series, plus a model that estimates the uplift from each trade promotion.

Models were trained in notebooks on laptops and refreshed once a quarter. Output went to the planning team as a spreadsheet. When the senior data scientist left, nobody could reproduce the last model: the training data had been a local extract and the code had changed since. Planners overrode about 60% of the forecasts, which told its own story.

Meanwhile the sales team was spending about 9% of revenue on trade promotions with little evidence of which schemes paid back. The COO wanted forecasts the planners trusted, a way to test promotions before running them, and a team that spends its time improving models instead of rebuilding them.

What could not be compromised

  • Every forecast in use must be traceable to the exact data, code and parameters that produced it.
  • No model goes live without a backtest report and sign-off from the demand planning head.
  • The ERP and planning systems stay the systems of record. The platform feeds them; it does not replace them.
  • A team of six data scientists and two data engineers must be able to run it, without a separate platform team.
  • Sales and pricing data stays within company-controlled infrastructure.
Options weighed

Four ways to run models in production

Options were compared on traceability, total cost over three years, fit with the team’s skills, and how much the forecasting approach could be shaped to the business.

OptionWhat worksWhat does notVerdict
Keep notebooks, add scheduled scriptsNo new tools, almost no cost.No tracking, no registry, no monitoring. The same fragility, now on a timer.Rejected
Demand-planning suite with built-in forecastingPackaged workflows for planners.Black-box models. Cannot add the company’s own promotion uplift features or distributor signals.Rejected
Managed cloud ML platformFast to start, little to operate.Usage costs grow with every retrain, sales data moves outside, and pipelines are tied to one provider.Kept as fallback
Open-source MLOps on Kubernetes: feature store, MLflow, pipelines, monitoringFull traceability, portable, fits the team’s Python skills, predictable cost.The company runs the platform, so it must be kept simple and automated.Chosen
Target architecture

One loop from data to forecast and back again

Data flows in from the ERP, distributors, promotions and outside sources into a feature store. Pipelines train and track candidate models, the best one is registered and sent for sign-off, CI/CD deploys it, and monitoring compares predictions with actual sales. When accuracy slips or the data shifts, the loop starts again on its own.

Scroll sideways to see the whole diagram →
BUSINESS DATA AND APPSMLOPS PLATFORM, ON KUBERNETESSHARED INFRASTRUCTUREPEOPLEERP and DMS salesprimary, secondary, dailyTrade promotionsschemes, spends, pricesExternal signalsweather, festivals, CPIPlanning systemsdemand plan, what-ifFeature storesales, promo, price featuresRetraining triggerdrift, schedule or new dataModel monitoringdata drift, forecast errorModel servingweekly batch, real-time APITraining pipelinesper SKU group, on KubernetesExperiment trackingMLflow runs, metrics, paramsModel registryMLflow, versions and stagesCI/CD for modelstests, backtests, canaryKubernetes cluster12 CPU + 2 GPU workersObject storageS3: datasets, artefactsModel sign-offdemand planning headData science team6 people, code in Gitruns, metricsbest candidateapproved versionpredictions and actualsdrift alert123456Data / replicationLogging / managementControl / API callException / alert
Numbered flows: (1) daily sales, promotion and external data load into the feature store, (2) training pipelines build point-in-time training sets and train candidates, (3) the best candidate is registered and the model owner approves it, (4) CI/CD deploys the approved version to serving, (5) forecasts and what-if results go to planning systems, (6) monitoring detects drift or rising error and triggers retraining.
Building blockWhy it is there
1 Feature storeOne definition for each feature, such as 4-week rolling sales, promotion depth, price index and festival flags, shared by training and serving. Point-in-time joins stop future data leaking into training. Built with Feast on PostgreSQL and object storage.
2 Training pipelinesArgo Workflows on Kubernetes train one gradient-boosted model per category group plus a hierarchical reconciliation step, so SKU, region and national forecasts add up. Runs are parallel, so a full retrain of all 50,000 series fits in under two hours.
3 Experiment tracking and model registryMLflow records every run with its data snapshot, code commit, parameters and metrics. The registry holds versions with stages: candidate, approved, production, retired. Nothing reaches production except through a registry stage change.
4 Human sign-offEach candidate arrives with a backtest report: error by category and region, bias, and the SKUs where it is worse than the current model. The demand planning head approves or rejects in the registry. Approval is recorded with name and date.
5 CI/CD for modelsPipeline code lives in Git. Every change runs unit tests, data-schema checks and a backtest against the current production model. Approved models deploy as a canary to five regions for one week before going national.
6 Batch and real-time servingA weekly batch writes 13-week forecasts for every SKU and region into the planning system. A small real-time API, served with KServe, answers promotion what-if questions from planners and key account managers in under 300 ms.
7 Monitoring and retraining triggersFeature drift is measured with population stability scores, and forecast error is tracked weekly against actual sales. A retrain starts on schedule every four weeks, or sooner when error rises or a feature drifts past its threshold.
Sizing, worked out

Sized from series count and retrain frequency, mostly on CPU

Gradient-boosted models on tabular data train well on CPU. GPUs are kept for a small amount of neural forecasting experimentation, not for production.

ItemFigureBasis
Forecast series~50,0002,000 SKUs x 25 regions, weekly grain
Training rows~7.8 million50,000 series x 156 weeks of history, about 180 features each
Full retrainUnder 2 hours12 category groups trained in parallel, ~40 minutes each on 32 cores, plus reconciliation and backtest
Compute12 CPU workers (64 cores, 512 GB) + 2 GPU workers (2 x L40S)Retrains, backtests and hyperparameter search at 3 to 4 parallel jobs; GPUs for experiments only
Weekly batch scoring~650,000 forecasts in ~15 minutes50,000 series x 13-week horizon on 4 workers
Real-time what-if APIPeak ~5 requests/second, p95 under 300 ms~150 planners and key account managers, 3 replicas for availability
Storage~8 TB usable, growing ~2 TB a yearFeature history, data snapshots for every run, model artefacts kept 3 years

The cluster can run on the company’s existing private cloud or virtualisation estate. Platform cost, including hardware over five years and support, is estimated at ₹55 to 80 lakh a year, depending on what infrastructure is already in place.

How it is delivered

Rebuild one category end to end, then scale the pattern

The first category goes through the whole loop, including sign-off and drift-triggered retraining, before any others move. That proves the platform on real decisions.

1

Foundations

Weeks 1 to 5

Kubernetes namespaces, object storage, MLflow and feature store set up. Data contracts agreed for ERP, DMS and promotion feeds.

Gate: Daily loads landing with schema checks passing for two weeks.

2

First category

Weeks 6 to 11

Home care forecasts rebuilt as pipelines with features, tracking and backtests. Current model reproduced as the baseline.

Gate: New model beats baseline error on 12 weeks of backtest.

3

Sign-off and serving

Weeks 12 to 15

Registry stages and approval flow live. Batch forecasts written to the planning system in parallel with the old spreadsheet.

Gate: Planners compare both for four weeks and approve the switch.

4

All categories

Weeks 16 to 21

Remaining 11 category groups move. Promotion uplift model and what-if API go live.

Gate: Every production forecast traceable to a registry version.

5

Monitoring and auto-retrain

Weeks 22 to 24

Drift and error thresholds tuned on live data. Scheduled and triggered retraining switched on.

Gate: One full retrain cycle completed and approved without manual pipeline work.

Way back: Every production model has a previous approved version in the registry. Rolling back is a stage change that redeploys the last version in minutes, and the weekly batch can be re-run from it. During migration, the old spreadsheet process runs in parallel until planners sign off.
Risks, handled up front

The risks that matter with forecasting models in production

RiskWhat could happenHow the design handles it
Data feed qualityA distributor stops sending data or sends duplicatesSchema and volume checks on every load. Missing regions are flagged, and forecasts for them are held, not filled with zeros.
Silent accuracy decayConsumer behaviour shifts after a price change or new competitorWeekly error tracking by category and region, drift scores on key features, and automatic retraining with sign-off.
Automation without judgementA retrained model goes live that is worse for key SKUsNo model reaches production without the backtest report and a named approval. Canary in five regions first.
Promotion effects misreadThe uplift model credits a scheme for sales that would have happened anywayBaselines built from non-promoted weeks and control regions. Uplift estimates come with ranges, not single numbers.
Platform burden on a small teamSix data scientists end up running infrastructureManaged by Git and templates, upgrades scheduled twice a year, and a support contract for Kubernetes and storage.
What was optimised

What was optimised

1 definition

Features shared by training and serving

The same feature code builds training sets and live inputs, so models see in production what they saw in training.

2 hours

Full retrain of 50,000 series

Parallel pipelines by category group replace a week of manual notebook work.

CPU first

Right hardware for tabular models

Gradient-boosted models run on CPU workers. Two GPU nodes are kept for experiments only.

4 weeks

Retrain cycle, or sooner

Scheduled retraining plus triggers on drift and error, instead of once a quarter.

1 click

Rollback to the last approved model

The registry keeps every approved version ready to redeploy.

100%

Forecasts traceable

Every number in the planning system links to data, code, parameters and an approver.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
Forecast error (weighted MAPE, SKU x region)~38%25 to 30%
Forecasts overridden by planners~60%Below 25%
Model refreshQuarterly, manualEvery 4 weeks or on drift, automatic with sign-off
Idea to production for a new model~3 months2 to 3 weeks
Inventory daysBaseline8 to 12% lower from better forecasts
Promotion evaluationAnecdotalEvery scheme scored for uplift before and after it runs

Targets are confirmed on the first category during backtesting. Accuracy gains depend on data quality from distributors and on how much of the current error comes from data rather than models.

Skills this draws on

What a team needs to deliver this

MLOps platform design

Feature stores, MLflow tracking and registry, pipelines and serving built as one loop on Kubernetes.

Demand forecasting

Hierarchical forecasts at SKU, region and national level, with reconciliation and honest backtesting.

Trade promotion analytics

Uplift models with proper baselines, so schemes are judged on incremental sales.

Model governance

Registry stages, backtest reports and approval flows that business owners actually use.

Monitoring and drift

Drift scores, error tracking and retraining triggers tuned to business calendars.

Kubernetes and data engineering

Reliable data feeds, schema contracts and platforms a small team can run.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

Do your models get rebuilt every quarter, by hand?

Tell us which models you run, where the data comes from and how forecasts reach the business. We will come back with a plain view of what an MLOps loop would look like for your team, and what it would take to get the first model through it.