Home / Deployment scenarios / SOC infrastructure across DC and DR
Reference deployment Security operations and DR

A bank’s own Security Operations Centre, built to keep watching when a data centre fails

A private bank with about 400 branches runs its security monitoring on an ageing log platform that fills up in six weeks and lives in a single data centre. This design replaces it with dedicated SOC infrastructure across the primary and DR sites, sized from real event rates, so the bank can search 180 days of logs, report incidents inside CERT-In’s six hours, and still see attacks on the day it loses its primary DC.

SectorPrivate bank, ~400 branches
Event rate~25,000 EPS average, 60,000 peak
Programme9 months, 5 phases
ModelOn-premises, primary plus warm DR
The situation

A log platform that forgets too soon and sits in one building

The bank runs core banking, UPI and card switching, internet and mobile banking, and around 400 branches with their own servers, firewalls and ATMs. Security logs from all of this flow into a log platform bought eight years ago. It ingests about 15,000 events per second on a good day, drops events at month end, and keeps only about six weeks of searchable data before it overwrites.

Two things forced the issue. CERT-In directions require logs of ICT systems to be kept for a rolling 180 days within India and incidents to be reported within six hours of being noticed. An RBI inspection also asked a pointed question: if the primary data centre is lost in an incident, who is watching the DR site while the bank recovers? The honest answer was nobody.

The CISO wanted a SOC the bank owns and runs: its own analysts, its own data, a platform that can grow for five years, and a DR posture that matches the rest of the bank. The brief was to design the infrastructure underneath, size it from real numbers, and make sure the SOC survives the same disasters the bank plans for.

What could not be compromised

  • Security logs and incident data stay inside the bank’s own data centres in India. No shared, multi-tenant platform.
  • Searchable retention of at least 180 days for every in-scope system, with a longer immutable archive under the bank’s retention policy.
  • Incidents must be reportable to CERT-In within six hours and to RBI as its cyber security framework expects, with evidence attached.
  • If the primary DC is lost, monitoring must resume from the DR site within 30 minutes, not after the bank has recovered.
  • Branch WAN links are often 4 to 10 Mbps. Log forwarding cannot compete with banking traffic or lose events when a link drops.
Options weighed

Four ways to run the SOC, compared honestly

Each option was scored on retention, survival of a site loss, control over data and five-year cost, using the bank’s own event rates.

OptionWhat worksWhat does notVerdict
Fully managed SOC on a provider’s shared platformFastest to start. Analysts come with the service.Logs leave the bank. Shared analysts, less control during an inspection, and fees grow with every new log source.Rejected
Single-site SOC, backups copied to DRCheapest hardware. Simple to run.If the primary DC goes, the bank is blind for days while the SIEM is rebuilt, which is exactly when attackers act.Rejected
Active-active SIEM stretched across both sitesNo failover step at all.Double hardware and licences, cross-site cluster traffic on every search, and a split-brain risk the team would have to manage.Rejected
Primary SOC with a warm DR SIEM fed in parallelDR sees every event in real time and keeps 30 days hot. Takeover is a switch, not a rebuild.About 35% more infrastructure than single-site. Two sets of content to keep in step.Chosen
Target architecture

One collection tier, three storage tiers, and a second SIEM that is already watching

Branches, data centre systems and security sensors send logs to one collection tier that parses, filters and buffers. A Kafka event bus decouples collection from the SIEM, so a slow search never drops an event. Data ages from fast NVMe to searchable erasure-coded disk to an immutable archive. The collectors send a full copy of every event to the DR site, where a smaller SIEM indexes it independently.

Scroll sideways to see the whole diagram →
OUTSIDE THE BANKLOG SOURCESPRIMARY DC: SOC PLATFORMDR DC: STANDBY SOCSOC TEAM, 24X7REGULATORSThreat intel feedsadvisories, sector sharingBranches400 log relaysDC systemscore banking, UPIEDR and NDRendpoints, tapsThreat intel platformIOC scoring, enrichmentCollection and parsing tier6 collectors, 60,000 EPS peak, store-and-forwardSIEM correlationrules, UEBA, ~400 use casesEvent busKafka, 72-hour replayHot tier: 30 daysNVMe indexers, 10 nodesSOARplaybooks, case managementCold archiveobject lock, kept 5 yearsWarm tier: to day 180searchable, erasure codedDR collectorsfull copy of every eventDR SIEM30 days hot, takes overArchive copyobject lock, second siteSOC analystsL1 to L3, three shiftsIncident responsetriage, contain, reportCERT-In and RBIincident reports, drill evidence1syslog over WANIOCsenrichparsedday 314alertscases6within 6 hoursindexsecond copy235Data / replicationControl / API callScheduled copyException / alert
Numbered flows: (1) branches and systems forward logs to the collection tier, (2) the event bus streams events to correlation, (3) and to the hot tier for indexing, (4) correlation raises alerts into SOAR and the analysts’ queue, (5) every event is also sent to the DR collectors in real time, (6) confirmed incidents are reported to CERT-In within six hours and to RBI as required. The archive is copied nightly to the DR site.
Building blockWhy it is there
1 Branch log relaysA small relay in each branch, often a virtual machine on the existing branch server, collects firewall, ATM controller, Active Directory and endpoint logs, compresses them and forwards them over the WAN with a rate cap. It holds 24 hours of logs locally if the link drops.
2 Collection and parsing tierSix collectors in the primary DC receive branch and data centre logs, normalise them to one schema, drop known noise and tag events with asset and user context. Each collector also forwards a full copy to the DR site.
3 Event busA three-broker Kafka cluster between collection and the SIEM. It keeps 72 hours of events, so the SIEM can be patched or rebuilt and catch up without losing anything, and new tools can read the same stream.
4 SIEM hot, warm and cold tiersTen NVMe indexers hold 30 days for fast search and correlation. A warm tier on high-capacity disks with erasure coding keeps days 31 to 180 searchable. An object store with object lock keeps compressed raw logs for five years, which no administrator can delete early.
5 Correlation and threat intelDetection rules, user and entity behaviour analytics and about 400 use cases mapped to attack techniques. A threat intel platform scores indicators from CERT-In advisories, sector sharing groups and commercial feeds before they reach the rules, to keep false positives down.
6 SOARPlaybooks for the twenty most common alert types: enrich, check, contain where safe (for example, isolate an endpoint through EDR or block an IP on the perimeter), and open a case. Every action is logged for the inspection file.
7 NDR and EDR feedsNetwork detection sensors on taps at the internet edge, the core banking segment and the DC-DR link, and EDR on servers and branch endpoints. Both send alerts, not raw packets, into the collection tier.
8 DR SIEM and analyst seatsA smaller SIEM at the DR site indexes the full live stream with 30 days hot, runs the same detection content, and has eight analyst seats ready. Older data is restored from the archive copy when needed.
Sizing, worked out

From events per second to terabytes, step by step

Event rates were measured for four weeks across a sample of 40 branches and every data centre source, then scaled. The average is about 25,000 events per second, with month-end and salary-day peaks reaching 60,000. The average event, after parsing, is about 600 bytes.

ItemFigureBasis
Daily volume~2.2 billion events, ~1.3 TB raw25,000 EPS x 86,400 seconds = 2.16 billion; x ~600 bytes each
Hot tier (30 days)~39 TB used, 48 TB usable NVMeIndexed and compressed to ~50% of raw = 0.65 TB/day x 30 days x 2 copies
Warm tier (days 31 to 180)~98 TB usable, ~160 TB raw disk0.65 TB/day x 150 days, erasure coded 4+2 (1.5x), plus 10% free space
Cold archive (5 years)~300 TB at each siteRaw compressed about 8:1 = ~0.16 TB/day x 365 x 5 years
Event bus buffer~12 TB across 3 brokers1.3 TB/day x 3 days x replication factor 3
Indexing capacity10 indexers, ~8,000 EPS each8 nodes cover the 60,000 EPS peak, plus 2 for failure and growth
DC to DR link for SOC traffic~300 Mbps peak, ~80 Mbps compressed60,000 EPS x 600 bytes = 36 MB/s, compressed about 4:1

Each branch averages about 15 events per second, roughly 70 kbps before compression, so log forwarding takes under 2% of a 4 Mbps branch link with a rate cap. Storage is sized for event growth of 25% a year for three years; indexers and disks are added as a purchase, not a redesign.

How it is delivered

Five phases, with the old platform running until the new one proves itself

The existing log platform keeps running in parallel until the new SOC has caught the same incidents for a full month, including one month end.

1

Assess and measure

Weeks 1 to 6

Log source inventory, four weeks of event-rate measurement, use-case list mapped to attack techniques and to RBI and CERT-In obligations. Hardware ordered in week 5.

Gate: Every in-scope source has an owner, an expected EPS and a parsing plan.

2

Build primary platform

Weeks 7 to 16

Collectors, event bus, SIEM tiers, SOAR and threat intel built from code in the primary DC. Data centre sources onboarded first.

Gate: Failure tests passed: indexer loss, broker loss, collector loss, with zero dropped events.

3

Branch rollout

Weeks 17 to 26

Relays deployed to branches in batches of 50, with rate caps tuned per link size. Old and new platforms both receive logs.

Gate: All 400 branches reporting, under 0.1% event loss measured end to end.

4

DR SIEM and content

Weeks 27 to 32

DR collectors, DR SIEM and archive replication go live. Detection content and playbooks are deployed to both sites from one repository.

Gate: A planted test incident is detected at both sites within the same minute.

5

Parallel run and takeover drill

Weeks 33 to 38

One month of parallel running, then a full takeover drill: the primary SIEM is isolated and analysts work from DR for a whole shift.

Gate: Monitoring resumes from DR in under 30 minutes, CISO and audit sign-off, old platform retired.

Way back: The old log platform keeps receiving the same feeds until the parallel run ends, so any gap in the new SOC can be covered by the old one. Branch relays can be pointed back with one configuration change pushed centrally.
Risks, handled up front

What could go wrong, and what is already in the plan

RiskWhat could happenHow the design handles it
Event volume grows faster than plannedNew sources such as cloud workloads double ingestion and the hot tier fills earlyFiltering at the collectors removes noise before it is indexed. Every new source has an EPS estimate before onboarding, and storage headroom is reviewed monthly.
Alert fatigueAnalysts drown in false positives and miss the real attackUse cases go live one at a time with tuning, threat intel is scored before it reaches rules, and SOAR handles the routine enrichment.
DR SIEM drifts from primaryRules or parsers differ, so DR misses what primary would catchContent is deployed to both sites from one Git repository. A weekly test event checks both sites detect the same thing.
Branch link saturationLog forwarding slows teller transactionsRate caps on each relay, compression, and 24 hours of local buffering. Banking traffic keeps priority on the WAN.
Attacker deletes evidenceAn insider or intruder with admin rights tries to erase logsObject lock on the archive, separate admin accounts for SOC platforms, and copies at both sites under different administrators.
What was optimised

Where the design saves money and effort

~35%

Noise filtered at collection

Duplicate and low-value events are dropped or summarised before indexing, which saves NVMe and licence volume.

3 tiers

Storage matched to age

Only 30 days sit on NVMe. Days 31 to 180 live on erasure-coded disk at a fraction of the cost per terabyte.

72 h

Replay instead of loss

The event bus lets the SIEM be patched or recovered without dropping a single event.

1 repo

Same content at both sites

Rules, parsers and playbooks are deployed from one repository, so DR is never a stale copy.

20

Playbooks for routine alerts

SOAR handles enrichment and safe containment for the most common alerts, so analysts focus on real investigations.

30 min

Takeover, not rebuild

DR already has every event indexed. Failing over means moving analysts, not restoring a SIEM.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
Searchable log retention~6 weeks180 days searchable, 5 years immutable
Peak ingestion without drops~15,000 EPS, drops at month end60,000 EPS with headroom, zero loss through the event bus
Monitoring after primary DC lossNone until rebuiltResumes from DR within 30 minutes
Time to detect a known attack patternHours to daysUnder 15 minutes for mapped use cases
CERT-In report preparationManual, about a dayDraft report from the case in under 2 hours, inside the 6-hour window
Inspection evidenceCollected by handCase history, playbook actions and drill results on record

Targets are confirmed during the measurement phase and the parallel run. Storage figures depend on the actual event mix, compression achieved and the number of log sources brought into scope.

Skills this draws on

What a team needs to deliver this

SIEM architecture

Hot, warm and cold tier design, indexer sizing and retention planning from measured event rates.

Log collection at scale

Branch relays, rate limiting, store-and-forward and parsing for hundreds of sites over thin WAN links.

Event streaming

Kafka design for buffering, replay and feeding several consumers from one stream.

Detection engineering

Use cases mapped to attack techniques, threat intel scoring and SOAR playbook design.

DC-DR architecture

Making security operations survive the same site failures the rest of the bank plans for.

Banking regulation

Translating CERT-In directions and RBI cyber security expectations into retention, reporting and drill evidence.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

Is your SOC sized for your real event rate, and would it survive losing your primary DC?

Share your log source list and a rough event count. We will come back with a plain sizing from EPS to terabytes, what retention really costs, and what it takes to keep watching from your DR site.