A bank’s own Security Operations Centre, built to keep watching when a data centre fails
A private bank with about 400 branches runs its security monitoring on an ageing log platform that fills up in six weeks and lives in a single data centre. This design replaces it with dedicated SOC infrastructure across the primary and DR sites, sized from real event rates, so the bank can search 180 days of logs, report incidents inside CERT-In’s six hours, and still see attacks on the day it loses its primary DC.
A log platform that forgets too soon and sits in one building
The bank runs core banking, UPI and card switching, internet and mobile banking, and around 400 branches with their own servers, firewalls and ATMs. Security logs from all of this flow into a log platform bought eight years ago. It ingests about 15,000 events per second on a good day, drops events at month end, and keeps only about six weeks of searchable data before it overwrites.
Two things forced the issue. CERT-In directions require logs of ICT systems to be kept for a rolling 180 days within India and incidents to be reported within six hours of being noticed. An RBI inspection also asked a pointed question: if the primary data centre is lost in an incident, who is watching the DR site while the bank recovers? The honest answer was nobody.
The CISO wanted a SOC the bank owns and runs: its own analysts, its own data, a platform that can grow for five years, and a DR posture that matches the rest of the bank. The brief was to design the infrastructure underneath, size it from real numbers, and make sure the SOC survives the same disasters the bank plans for.
What could not be compromised
- Security logs and incident data stay inside the bank’s own data centres in India. No shared, multi-tenant platform.
- Searchable retention of at least 180 days for every in-scope system, with a longer immutable archive under the bank’s retention policy.
- Incidents must be reportable to CERT-In within six hours and to RBI as its cyber security framework expects, with evidence attached.
- If the primary DC is lost, monitoring must resume from the DR site within 30 minutes, not after the bank has recovered.
- Branch WAN links are often 4 to 10 Mbps. Log forwarding cannot compete with banking traffic or lose events when a link drops.
Four ways to run the SOC, compared honestly
Each option was scored on retention, survival of a site loss, control over data and five-year cost, using the bank’s own event rates.
One collection tier, three storage tiers, and a second SIEM that is already watching
Branches, data centre systems and security sensors send logs to one collection tier that parses, filters and buffers. A Kafka event bus decouples collection from the SIEM, so a slow search never drops an event. Data ages from fast NVMe to searchable erasure-coded disk to an immutable archive. The collectors send a full copy of every event to the DR site, where a smaller SIEM indexes it independently.
| Building block | Why it is there |
|---|---|
| 1 Branch log relays | A small relay in each branch, often a virtual machine on the existing branch server, collects firewall, ATM controller, Active Directory and endpoint logs, compresses them and forwards them over the WAN with a rate cap. It holds 24 hours of logs locally if the link drops. |
| 2 Collection and parsing tier | Six collectors in the primary DC receive branch and data centre logs, normalise them to one schema, drop known noise and tag events with asset and user context. Each collector also forwards a full copy to the DR site. |
| 3 Event bus | A three-broker Kafka cluster between collection and the SIEM. It keeps 72 hours of events, so the SIEM can be patched or rebuilt and catch up without losing anything, and new tools can read the same stream. |
| 4 SIEM hot, warm and cold tiers | Ten NVMe indexers hold 30 days for fast search and correlation. A warm tier on high-capacity disks with erasure coding keeps days 31 to 180 searchable. An object store with object lock keeps compressed raw logs for five years, which no administrator can delete early. |
| 5 Correlation and threat intel | Detection rules, user and entity behaviour analytics and about 400 use cases mapped to attack techniques. A threat intel platform scores indicators from CERT-In advisories, sector sharing groups and commercial feeds before they reach the rules, to keep false positives down. |
| 6 SOAR | Playbooks for the twenty most common alert types: enrich, check, contain where safe (for example, isolate an endpoint through EDR or block an IP on the perimeter), and open a case. Every action is logged for the inspection file. |
| 7 NDR and EDR feeds | Network detection sensors on taps at the internet edge, the core banking segment and the DC-DR link, and EDR on servers and branch endpoints. Both send alerts, not raw packets, into the collection tier. |
| 8 DR SIEM and analyst seats | A smaller SIEM at the DR site indexes the full live stream with 30 days hot, runs the same detection content, and has eight analyst seats ready. Older data is restored from the archive copy when needed. |
From events per second to terabytes, step by step
Event rates were measured for four weeks across a sample of 40 branches and every data centre source, then scaled. The average is about 25,000 events per second, with month-end and salary-day peaks reaching 60,000. The average event, after parsing, is about 600 bytes.
| Item | Figure | Basis |
|---|---|---|
| Daily volume | ~2.2 billion events, ~1.3 TB raw | 25,000 EPS x 86,400 seconds = 2.16 billion; x ~600 bytes each |
| Hot tier (30 days) | ~39 TB used, 48 TB usable NVMe | Indexed and compressed to ~50% of raw = 0.65 TB/day x 30 days x 2 copies |
| Warm tier (days 31 to 180) | ~98 TB usable, ~160 TB raw disk | 0.65 TB/day x 150 days, erasure coded 4+2 (1.5x), plus 10% free space |
| Cold archive (5 years) | ~300 TB at each site | Raw compressed about 8:1 = ~0.16 TB/day x 365 x 5 years |
| Event bus buffer | ~12 TB across 3 brokers | 1.3 TB/day x 3 days x replication factor 3 |
| Indexing capacity | 10 indexers, ~8,000 EPS each | 8 nodes cover the 60,000 EPS peak, plus 2 for failure and growth |
| DC to DR link for SOC traffic | ~300 Mbps peak, ~80 Mbps compressed | 60,000 EPS x 600 bytes = 36 MB/s, compressed about 4:1 |
Each branch averages about 15 events per second, roughly 70 kbps before compression, so log forwarding takes under 2% of a 4 Mbps branch link with a rate cap. Storage is sized for event growth of 25% a year for three years; indexers and disks are added as a purchase, not a redesign.
Five phases, with the old platform running until the new one proves itself
The existing log platform keeps running in parallel until the new SOC has caught the same incidents for a full month, including one month end.
Assess and measure
Weeks 1 to 6
Log source inventory, four weeks of event-rate measurement, use-case list mapped to attack techniques and to RBI and CERT-In obligations. Hardware ordered in week 5.
Gate: Every in-scope source has an owner, an expected EPS and a parsing plan.
Build primary platform
Weeks 7 to 16
Collectors, event bus, SIEM tiers, SOAR and threat intel built from code in the primary DC. Data centre sources onboarded first.
Gate: Failure tests passed: indexer loss, broker loss, collector loss, with zero dropped events.
Branch rollout
Weeks 17 to 26
Relays deployed to branches in batches of 50, with rate caps tuned per link size. Old and new platforms both receive logs.
Gate: All 400 branches reporting, under 0.1% event loss measured end to end.
DR SIEM and content
Weeks 27 to 32
DR collectors, DR SIEM and archive replication go live. Detection content and playbooks are deployed to both sites from one repository.
Gate: A planted test incident is detected at both sites within the same minute.
Parallel run and takeover drill
Weeks 33 to 38
One month of parallel running, then a full takeover drill: the primary SIEM is isolated and analysts work from DR for a whole shift.
Gate: Monitoring resumes from DR in under 30 minutes, CISO and audit sign-off, old platform retired.
What could go wrong, and what is already in the plan
| Risk | What could happen | How the design handles it |
|---|---|---|
| Event volume grows faster than planned | New sources such as cloud workloads double ingestion and the hot tier fills early | Filtering at the collectors removes noise before it is indexed. Every new source has an EPS estimate before onboarding, and storage headroom is reviewed monthly. |
| Alert fatigue | Analysts drown in false positives and miss the real attack | Use cases go live one at a time with tuning, threat intel is scored before it reaches rules, and SOAR handles the routine enrichment. |
| DR SIEM drifts from primary | Rules or parsers differ, so DR misses what primary would catch | Content is deployed to both sites from one Git repository. A weekly test event checks both sites detect the same thing. |
| Branch link saturation | Log forwarding slows teller transactions | Rate caps on each relay, compression, and 24 hours of local buffering. Banking traffic keeps priority on the WAN. |
| Attacker deletes evidence | An insider or intruder with admin rights tries to erase logs | Object lock on the archive, separate admin accounts for SOC platforms, and copies at both sites under different administrators. |
Where the design saves money and effort
Noise filtered at collection
Duplicate and low-value events are dropped or summarised before indexing, which saves NVMe and licence volume.
Storage matched to age
Only 30 days sit on NVMe. Days 31 to 180 live on erasure-coded disk at a fraction of the cost per terabyte.
Replay instead of loss
The event bus lets the SIEM be patched or recovered without dropping a single event.
Same content at both sites
Rules, parsers and playbooks are deployed from one repository, so DR is never a stale copy.
Playbooks for routine alerts
SOAR handles enrichment and safe containment for the most common alerts, so analysts focus on real investigations.
Takeover, not rebuild
DR already has every event indexed. Failing over means moving analysts, not restoring a SIEM.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Searchable log retention | ~6 weeks | 180 days searchable, 5 years immutable |
| Peak ingestion without drops | ~15,000 EPS, drops at month end | 60,000 EPS with headroom, zero loss through the event bus |
| Monitoring after primary DC loss | None until rebuilt | Resumes from DR within 30 minutes |
| Time to detect a known attack pattern | Hours to days | Under 15 minutes for mapped use cases |
| CERT-In report preparation | Manual, about a day | Draft report from the case in under 2 hours, inside the 6-hour window |
| Inspection evidence | Collected by hand | Case history, playbook actions and drill results on record |
Targets are confirmed during the measurement phase and the parallel run. Storage figures depend on the actual event mix, compression achieved and the number of log sources brought into scope.
What a team needs to deliver this
SIEM architecture
Hot, warm and cold tier design, indexer sizing and retention planning from measured event rates.
Log collection at scale
Branch relays, rate limiting, store-and-forward and parsing for hundreds of sites over thin WAN links.
Event streaming
Kafka design for buffering, replay and feeding several consumers from one stream.
Detection engineering
Use cases mapped to attack techniques, threat intel scoring and SOAR playbook design.
DC-DR architecture
Making security operations survive the same site failures the rest of the bank plans for.
Banking regulation
Translating CERT-In directions and RBI cyber security expectations into retention, reporting and drill evidence.
Other scenarios
Three-site core banking DR
A small finance bank protects its core banking and payment switch across three sites, with zero data loss to a near DR site and a far site in another seismic zone.
Read →Disaster recoveryMulti-tenant DR as a Service
A regional data-centre operator builds a shared DR service for 40 to 60 mid-size customers, with isolated tenant networks, self-service testing and billing per protected VM.
Read →Cyber recoveryCyber recovery vault for ransomware
A listed pharmaceutical company builds an isolated recovery vault and clean room, so it can rebuild after ransomware from copies attackers cannot reach.
Read →Is your SOC sized for your real event rate, and would it survive losing your primary DC?
Share your log source list and a rough event count. We will come back with a plain sizing from EPS to terabytes, what retention really costs, and what it takes to keep watching from your DR site.