Home / Deployment scenarios / Three-site core banking DR
Reference deployment Disaster recovery

Core banking that loses no transactions, even when a whole city has a bad day

A small finance bank runs core banking, its payment switch and UPI from one data centre, with a DR site 1,400 km away that trails by up to fifteen minutes. For a bank posting thousands of payments a minute, fifteen minutes of lost data is not a recovery, it is a reconciliation crisis. This design adds a near DR site in the same metro with synchronous replication, keeps the far site for regional disasters, and makes switching between them something the bank rehearses, not something it hopes for.

SectorSmall finance bank
Scale~250 branches, ~3 million accounts
SitesPrimary, near DR, far DR
TargetsRPO 0 near, under 5 min far
The situation

A DR site that is safe but always a few minutes behind

The bank serves about three million customers through some 250 branches, a mobile app, internet banking and a growing UPI business. Core banking, the payment switch, ATM and card processing, and the interfaces to UPI, IMPS, NEFT and RTGS all run in a primary data centre in a metro city. A DR site in another region receives asynchronous database and storage replication.

In practice the DR site lags by five to fifteen minutes, more during end-of-day batch. A power incident at the primary DC last year came close to forcing a failover. The post-incident review was uncomfortable: thousands of UPI and IMPS transactions would have been in flight, customers would have been debited with no record at DR, and reconciliation with the payment networks would have taken days.

The board asked for zero data loss for the systems that move money, without giving up protection against a regional event such as a flood or an earthquake. RBI expects regular DR drills for critical systems, run from the DR site for a real business period, and the payment network operator expects members to prove they can switch sites and come back. Any design had to make those drills routine.

What could not be compromised

  • Zero data loss for core banking and payments in any single-site failure. Synchronous replication cannot add more than about 1 ms to a commit.
  • Protection against a regional event: one copy of the data at least a few hundred kilometres away, in a different seismic zone.
  • Payment switch, UPI, IMPS and card connectivity must be live and certified at the DR sites before they are needed, not set up during an incident.
  • DR drills for critical systems at least every six months, run from DR for a full business day, with evidence for the board and the regulator.
  • The core banking vendor supports one active database at a time. No active-active writes across sites.
Options weighed

Four ways to protect the core, compared honestly

Each option was assessed against the bank’s worst realistic scenarios: loss of the primary DC, loss of the metro, and a failure during end-of-day batch.

OptionWhat worksWhat does notVerdict
Stay with two sites, asynchronous to far DRAlready in place, protects against regional events.Minutes of lost transactions in every failover. The exact problem the board raised.Rejected
Two sites, synchronous to a near DR onlyZero data loss, simple.Both sites in the same metro and seismic zone. A regional flood or earthquake takes out everything.Rejected
Active-active core banking across two sitesNo failover step for the core.Not supported by the core banking product, and stretched clusters bring split-brain risk.Rejected
Three sites: synchronous near DR, asynchronous far DRZero data loss for site failures, regional protection kept, far DR can be fed from either site.A third site to fund and run, and a metro fibre pair with tight latency.Chosen
Target architecture

Synchronous within the metro, asynchronous across the country

Every committed transaction is written to the near DR database before the customer sees a success. The near site then cascades changes to the far site, so the primary does not wait on a long-distance link. A quorum observer at a third location decides whether a failover is safe, and the orchestrator runs the same runbook in a drill or a real event.

Scroll sideways to see the whole diagram →
CUSTOMERS AND PAYMENT NETWORKSASSURANCESITE CONTROLPRIMARY DCNEAR DR: SAME METROFAR DR: OTHER ZONECustomer channelsbranch, mobile, webPayment networksUPI, IMPS, NEFT, RTGS, cardsDrill recordsevidence per drillRegulatordrill, outage reportsTraffic managerDNS steering, 3 sitesQuorum observerdecides failoverDR orchestratorrunbooks, switchoverPayment switchactive, UPI, IMPS, cardsCBS application tieractive, 12 app serversCore databaseclustered, 2 nodes, ~6 TBAll-flash storagefiles, batch, documentsPayment switchwarm standby, certifiedCBS application tierwarm, 8 of 12 serversDatabase standbysync, zero data lossSYNCStorage replicasync over metro fibrePayment switchstandby, certifiedCBS application tier4 running, 8 from imagesDatabase standbyasync, lag under 5 minStorage replicaasync, journal-basedpostSQLsyncasync123456User or API trafficControl / API callData / replicationSynchronous writeLogging / management
Numbered flows: (1) channels reach core banking through a traffic manager that can steer to any site, (2) payment networks connect to the switch at the primary and the near DR site, (3) every commit is written synchronously to the near DR standby, (4) the near standby cascades changes asynchronously to the far site, (5) a quorum observer watches the primary and decides on automatic failover to near DR, (6) the orchestrator runs switchover runbooks for storage, database, applications and DNS.
Building blockWhy it is there
1 Core database with synchronous standbyThe primary database ships every change to a standby in the near DR site and waits for acknowledgement before a commit completes. A standby in the far site receives changes asynchronously, cascaded from the near site, and can be re-pointed to the primary if the near site is down.
2 Quorum observer at a third locationA lightweight observer hosted at the far DR site watches both the primary and the near standby. Automatic failover only happens when the observer and the standby agree the primary is gone, which prevents two active databases.
3 Storage replicationInterface files, batch outputs, statements and documents live on all-flash arrays. They replicate synchronously to the near site and asynchronously to the far site with consistent journals, so file and database recovery points line up.
4 CBS application tierTwelve application servers at the primary. Eight run warm at the near site with the same versions and configuration, deployed from the same pipeline. The far site keeps four running and builds the rest from golden images when invoked.
5 Payment switch and network linksThe switch runs active at the primary and warm at the near site, with links to UPI, IMPS, card and ATM networks live and certified at both. The far site has standby links that are tested in every drill.
6 Traffic manager and DNSBranch, mobile and internet channels reach the core through a global traffic manager. On switchover it steers new sessions to the active site within a minute, and branch WAN routers follow the same change.
7 DR orchestrator and runbooksScripted runbooks for planned switchover, unplanned failover and failback: storage role change, database role change, application start order, switch activation, DNS. The same runbook runs in a drill and in a real event.
Sizing, worked out

Recovery targets by tier, and the links that make them possible

Every application was placed in a tier with the business owner, then link capacity and latency were worked out from measured database change rates. The core posts about 250 financial transactions a second at peak, and end-of-day batch generates about 2.5 times the online change rate.

ItemNear DR (same metro)Far DR (other zone)Basis
Tier 1: core banking, payment switch, UPI, ATMRPO 0, RTO 30 minutesRPO under 5 minutes, RTO 2 hoursSynchronous database and storage to near, cascaded to far, warm app tier at near
Tier 2: internet and mobile banking, treasury, loan originationRPO under 5 minutes, RTO 2 hoursRPO 15 minutes, RTO 4 hoursAsynchronous replication, started after Tier 1 in runbook order
Tier 3: MIS, HR, email, intranetNot protected at nearRPO 24 hours, RTO 24 to 48 hoursRestored from backups copied nightly to far DR
Database change rate~10 MB/s online, ~25 MB/s batchSame stream, buffered250 TPS x ~40 KB of change per posting; batch measured at 2.5x online
Replication bandwidth2 x 10G on diverse fibre routes2 x 1G from two carriers25 MB/s = 200 Mbps database + ~400 Mbps storage peak = ~0.6 Gbps; x4 headroom near, far lag held under 5 minutes
Distance and latency~35 km route, ~0.4 ms round trip~1,400 km, ~25 ms round tripLight in fibre takes about 5 microseconds per km each way, plus equipment. Synchronous is only viable near
Application capacity at DR8 of 12 servers warm, 12 in 20 minutes4 running, 12 in about 90 minutesPeak needs 9 servers; warm capacity covers normal load at once

The near DR site is chosen for fibre route length, not straight-line distance: two physically separate routes under 40 km each, on different power grids and flood plains from the primary. The far site is in a different seismic zone, more than 1,000 km away.

How it is delivered

Five phases, with the far DR protecting the bank the whole time

The existing far DR keeps running throughout. The near site is added alongside it, and the cascade is only switched on once the near site has proved itself.

1

Tiering and site selection

Weeks 1 to 6

Every application tiered with its owner, RPO and RTO agreed. Near DR site and fibre routes selected on route length, power and flood risk.

Gate: Board approves tiers and targets. Fibre round trip measured under 0.5 ms.

2

Build near DR

Weeks 7 to 18

Racks, network, storage, database standby and app tier built from the same code and images as the primary. Payment network links ordered and certified.

Gate: Synchronous replication running for 30 days with commit latency increase under 1 ms at peak.

3

Cascade and observer

Weeks 19 to 22

Far DR re-pointed to receive changes from the near site. Quorum observer deployed at the far site. Re-pointing far DR to the primary rehearsed.

Gate: Far lag under 5 minutes through three end-of-day batches.

4

Runbooks and first drill

Weeks 23 to 28

Orchestrated runbooks for switchover, failover and failback. First planned switchover to near DR on a weekend, then a full business day run from near DR.

Gate: Zero transactions lost, payment networks reconciled, RTO measured under 30 minutes.

5

Regulator-grade drills

Ongoing, every 6 months

Alternate drills: switch to near DR, and switch to far DR. Each run for a full business day, with evidence collected by the orchestrator.

Gate: Drill report to the board and regulator, issues closed before the next drill.

Way back: Until the cascade is switched on, the far DR keeps receiving changes directly from the primary exactly as it does today. If synchronous replication ever adds too much latency, it can be switched to asynchronous mode in seconds while the cause is fixed, and the bank is no worse off than before.
Risks, handled up front

What could go wrong, and what is already in the plan

RiskWhat could happenHow the design handles it
Metro link failureA fibre cut stalls synchronous commits and slows every transactionTwo physically diverse routes. If both fail, the database drops to asynchronous mode automatically after a short timeout and alerts the operations team.
Split brainBoth primary and near standby believe they are activeAutomatic failover requires agreement from the quorum observer at the far site. Manual failover requires two named approvers.
Near DR site downThe cascade stops and far DR falls behindFar DR is re-pointed to receive directly from the primary, a step rehearsed in every drill.
Payment networks not ready at DRCore banking fails over but UPI or IMPS stays downLinks live and certified at the near site, tested in every drill with real transactions, and switch activation is part of the runbook.
Configuration driftDR app servers run an older version or setting and fail on switchoverAll sites deployed from one pipeline and one set of images. A daily check compares versions and configuration across sites.
What was optimised

Where the design keeps cost and risk down

0

Data loss in a site failure

Synchronous replication to the near site, with the latency budget measured before the site was chosen.

1

Long-distance stream

The far site is fed by cascade from near DR, so the primary never waits on a 1,400 km link.

3 tiers

Protection matched to need

Only Tier 1 gets synchronous protection. Tier 3 systems are restored from backups at the far site.

8 of 12

Warm, not full, capacity

Near DR runs enough warm servers for normal load and scales to full within 20 minutes.

1 runbook

Same steps in a drill or a disaster

Switchover is scripted end to end, so a drill is a rehearsal of the real thing.

Every 6 months

Drills as routine

Alternating near and far drills, each run for a full business day with evidence captured automatically.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
Data loss in a primary site failure5 to 15 minutesZero for Tier 1
Recovery time for core banking and payments4 to 6 hours, mostly manualUnder 30 minutes to near DR, under 2 hours to far DR
Data loss in a regional disaster5 to 15 minutesUnder 5 minutes
Commit latency added-Under 1 ms at peak
DR drillsAnnual, partial, weekend onlyEvery 6 months, full business day, alternating sites
Reconciliation after failoverDays, with payment networksSame day, nothing missing to reconcile

Targets are confirmed during the build and the first drill. Latency figures depend on the actual fibre routes, and recovery times depend on the core banking product’s start-up behaviour, which is measured before targets are signed off.

Skills this draws on

What a team needs to deliver this

DC-DR architecture

Three-site designs, site selection on route length and hazard, and recovery tiers agreed with the business.

Database replication

Synchronous and asynchronous standbys, cascading, observers and safe automatic failover.

Storage replication

Consistent journals across arrays so file and database recovery points match.

Payment systems continuity

Switch and network connectivity at DR sites, certification and real-transaction drill testing.

Runbook automation

Scripted switchover and failback across storage, database, applications and DNS.

Banking regulation

DR drill programmes and evidence that stand up to board, audit and regulator review.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

How many transactions would your bank lose if the primary DC failed right now?

Share your current DR setup, distances between sites and a rough change rate. We will come back with a plain view of what near-zero data loss would take, what it would cost, and how to drill it without fear.