Core banking that loses no transactions, even when a whole city has a bad day
A small finance bank runs core banking, its payment switch and UPI from one data centre, with a DR site 1,400 km away that trails by up to fifteen minutes. For a bank posting thousands of payments a minute, fifteen minutes of lost data is not a recovery, it is a reconciliation crisis. This design adds a near DR site in the same metro with synchronous replication, keeps the far site for regional disasters, and makes switching between them something the bank rehearses, not something it hopes for.
A DR site that is safe but always a few minutes behind
The bank serves about three million customers through some 250 branches, a mobile app, internet banking and a growing UPI business. Core banking, the payment switch, ATM and card processing, and the interfaces to UPI, IMPS, NEFT and RTGS all run in a primary data centre in a metro city. A DR site in another region receives asynchronous database and storage replication.
In practice the DR site lags by five to fifteen minutes, more during end-of-day batch. A power incident at the primary DC last year came close to forcing a failover. The post-incident review was uncomfortable: thousands of UPI and IMPS transactions would have been in flight, customers would have been debited with no record at DR, and reconciliation with the payment networks would have taken days.
The board asked for zero data loss for the systems that move money, without giving up protection against a regional event such as a flood or an earthquake. RBI expects regular DR drills for critical systems, run from the DR site for a real business period, and the payment network operator expects members to prove they can switch sites and come back. Any design had to make those drills routine.
What could not be compromised
- Zero data loss for core banking and payments in any single-site failure. Synchronous replication cannot add more than about 1 ms to a commit.
- Protection against a regional event: one copy of the data at least a few hundred kilometres away, in a different seismic zone.
- Payment switch, UPI, IMPS and card connectivity must be live and certified at the DR sites before they are needed, not set up during an incident.
- DR drills for critical systems at least every six months, run from DR for a full business day, with evidence for the board and the regulator.
- The core banking vendor supports one active database at a time. No active-active writes across sites.
Four ways to protect the core, compared honestly
Each option was assessed against the bank’s worst realistic scenarios: loss of the primary DC, loss of the metro, and a failure during end-of-day batch.
Synchronous within the metro, asynchronous across the country
Every committed transaction is written to the near DR database before the customer sees a success. The near site then cascades changes to the far site, so the primary does not wait on a long-distance link. A quorum observer at a third location decides whether a failover is safe, and the orchestrator runs the same runbook in a drill or a real event.
| Building block | Why it is there |
|---|---|
| 1 Core database with synchronous standby | The primary database ships every change to a standby in the near DR site and waits for acknowledgement before a commit completes. A standby in the far site receives changes asynchronously, cascaded from the near site, and can be re-pointed to the primary if the near site is down. |
| 2 Quorum observer at a third location | A lightweight observer hosted at the far DR site watches both the primary and the near standby. Automatic failover only happens when the observer and the standby agree the primary is gone, which prevents two active databases. |
| 3 Storage replication | Interface files, batch outputs, statements and documents live on all-flash arrays. They replicate synchronously to the near site and asynchronously to the far site with consistent journals, so file and database recovery points line up. |
| 4 CBS application tier | Twelve application servers at the primary. Eight run warm at the near site with the same versions and configuration, deployed from the same pipeline. The far site keeps four running and builds the rest from golden images when invoked. |
| 5 Payment switch and network links | The switch runs active at the primary and warm at the near site, with links to UPI, IMPS, card and ATM networks live and certified at both. The far site has standby links that are tested in every drill. |
| 6 Traffic manager and DNS | Branch, mobile and internet channels reach the core through a global traffic manager. On switchover it steers new sessions to the active site within a minute, and branch WAN routers follow the same change. |
| 7 DR orchestrator and runbooks | Scripted runbooks for planned switchover, unplanned failover and failback: storage role change, database role change, application start order, switch activation, DNS. The same runbook runs in a drill and in a real event. |
Recovery targets by tier, and the links that make them possible
Every application was placed in a tier with the business owner, then link capacity and latency were worked out from measured database change rates. The core posts about 250 financial transactions a second at peak, and end-of-day batch generates about 2.5 times the online change rate.
| Item | Near DR (same metro) | Far DR (other zone) | Basis |
|---|---|---|---|
| Tier 1: core banking, payment switch, UPI, ATM | RPO 0, RTO 30 minutes | RPO under 5 minutes, RTO 2 hours | Synchronous database and storage to near, cascaded to far, warm app tier at near |
| Tier 2: internet and mobile banking, treasury, loan origination | RPO under 5 minutes, RTO 2 hours | RPO 15 minutes, RTO 4 hours | Asynchronous replication, started after Tier 1 in runbook order |
| Tier 3: MIS, HR, email, intranet | Not protected at near | RPO 24 hours, RTO 24 to 48 hours | Restored from backups copied nightly to far DR |
| Database change rate | ~10 MB/s online, ~25 MB/s batch | Same stream, buffered | 250 TPS x ~40 KB of change per posting; batch measured at 2.5x online |
| Replication bandwidth | 2 x 10G on diverse fibre routes | 2 x 1G from two carriers | 25 MB/s = 200 Mbps database + ~400 Mbps storage peak = ~0.6 Gbps; x4 headroom near, far lag held under 5 minutes |
| Distance and latency | ~35 km route, ~0.4 ms round trip | ~1,400 km, ~25 ms round trip | Light in fibre takes about 5 microseconds per km each way, plus equipment. Synchronous is only viable near |
| Application capacity at DR | 8 of 12 servers warm, 12 in 20 minutes | 4 running, 12 in about 90 minutes | Peak needs 9 servers; warm capacity covers normal load at once |
The near DR site is chosen for fibre route length, not straight-line distance: two physically separate routes under 40 km each, on different power grids and flood plains from the primary. The far site is in a different seismic zone, more than 1,000 km away.
Five phases, with the far DR protecting the bank the whole time
The existing far DR keeps running throughout. The near site is added alongside it, and the cascade is only switched on once the near site has proved itself.
Tiering and site selection
Weeks 1 to 6
Every application tiered with its owner, RPO and RTO agreed. Near DR site and fibre routes selected on route length, power and flood risk.
Gate: Board approves tiers and targets. Fibre round trip measured under 0.5 ms.
Build near DR
Weeks 7 to 18
Racks, network, storage, database standby and app tier built from the same code and images as the primary. Payment network links ordered and certified.
Gate: Synchronous replication running for 30 days with commit latency increase under 1 ms at peak.
Cascade and observer
Weeks 19 to 22
Far DR re-pointed to receive changes from the near site. Quorum observer deployed at the far site. Re-pointing far DR to the primary rehearsed.
Gate: Far lag under 5 minutes through three end-of-day batches.
Runbooks and first drill
Weeks 23 to 28
Orchestrated runbooks for switchover, failover and failback. First planned switchover to near DR on a weekend, then a full business day run from near DR.
Gate: Zero transactions lost, payment networks reconciled, RTO measured under 30 minutes.
Regulator-grade drills
Ongoing, every 6 months
Alternate drills: switch to near DR, and switch to far DR. Each run for a full business day, with evidence collected by the orchestrator.
Gate: Drill report to the board and regulator, issues closed before the next drill.
What could go wrong, and what is already in the plan
| Risk | What could happen | How the design handles it |
|---|---|---|
| Metro link failure | A fibre cut stalls synchronous commits and slows every transaction | Two physically diverse routes. If both fail, the database drops to asynchronous mode automatically after a short timeout and alerts the operations team. |
| Split brain | Both primary and near standby believe they are active | Automatic failover requires agreement from the quorum observer at the far site. Manual failover requires two named approvers. |
| Near DR site down | The cascade stops and far DR falls behind | Far DR is re-pointed to receive directly from the primary, a step rehearsed in every drill. |
| Payment networks not ready at DR | Core banking fails over but UPI or IMPS stays down | Links live and certified at the near site, tested in every drill with real transactions, and switch activation is part of the runbook. |
| Configuration drift | DR app servers run an older version or setting and fail on switchover | All sites deployed from one pipeline and one set of images. A daily check compares versions and configuration across sites. |
Where the design keeps cost and risk down
Data loss in a site failure
Synchronous replication to the near site, with the latency budget measured before the site was chosen.
Long-distance stream
The far site is fed by cascade from near DR, so the primary never waits on a 1,400 km link.
Protection matched to need
Only Tier 1 gets synchronous protection. Tier 3 systems are restored from backups at the far site.
Warm, not full, capacity
Near DR runs enough warm servers for normal load and scales to full within 20 minutes.
Same steps in a drill or a disaster
Switchover is scripted end to end, so a drill is a rehearsal of the real thing.
Drills as routine
Alternating near and far drills, each run for a full business day with evidence captured automatically.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Data loss in a primary site failure | 5 to 15 minutes | Zero for Tier 1 |
| Recovery time for core banking and payments | 4 to 6 hours, mostly manual | Under 30 minutes to near DR, under 2 hours to far DR |
| Data loss in a regional disaster | 5 to 15 minutes | Under 5 minutes |
| Commit latency added | - | Under 1 ms at peak |
| DR drills | Annual, partial, weekend only | Every 6 months, full business day, alternating sites |
| Reconciliation after failover | Days, with payment networks | Same day, nothing missing to reconcile |
Targets are confirmed during the build and the first drill. Latency figures depend on the actual fibre routes, and recovery times depend on the core banking product’s start-up behaviour, which is measured before targets are signed off.
What a team needs to deliver this
DC-DR architecture
Three-site designs, site selection on route length and hazard, and recovery tiers agreed with the business.
Database replication
Synchronous and asynchronous standbys, cascading, observers and safe automatic failover.
Storage replication
Consistent journals across arrays so file and database recovery points match.
Payment systems continuity
Switch and network connectivity at DR sites, certification and real-transaction drill testing.
Runbook automation
Scripted switchover and failback across storage, database, applications and DNS.
Banking regulation
DR drill programmes and evidence that stand up to board, audit and regulator review.
Other scenarios
Multi-tenant DR as a Service
A regional data-centre operator builds a shared DR service for 40 to 60 mid-size customers, with isolated tenant networks, self-service testing and billing per protected VM.
Read →Cyber recoveryCyber recovery vault for ransomware
A listed pharmaceutical company builds an isolated recovery vault and clean room, so it can rebuild after ransomware from copies attackers cannot reach.
Read →Security operations and DRSOC infrastructure across DC and DR
A private bank builds its own Security Operations Centre across two data centres, sized from 60,000 events per second, with a DR SIEM that can take over.
Read →How many transactions would your bank lose if the primary DC failed right now?
Share your current DR setup, distances between sites and a rough change rate. We will come back with a plain view of what near-zero data loss would take, what it would cost, and how to drill it without fear.