A private AI platform with no connection to the outside world
A public research laboratory holds decades of technical reports that fall under export control. Its scientists want the same AI help their peers elsewhere take for granted, but the data cannot sit on any network that touches the internet. This design builds a complete AI platform inside the laboratory’s isolated network, and treats every model, package and patch that crosses the air gap as something to be checked, signed and recorded.
Valuable knowledge, locked in reports nobody can search
The laboratory has about 350 scientists and engineers cleared to work on its classified network. Over thirty years they have written roughly 1.2 lakh technical reports, test records and design notes. Much of it is export-controlled technical data under the SCOMET list, and all of it lives on a network with no physical connection to the internet.
Finding prior work is slow. Search on the isolated network is a basic keyword tool over file shares, and senior staff estimate that engineers spend a day or more a week looking for earlier results, or repeating tests because they could not find them. Several research groups also want to fine-tune models on their own data, which today means writing code on paper notes and carrying it in.
Public AI services are out of the question, and so is a normal on-premises platform that pulls updates from the internet. The director’s brief is direct: give researchers a capable assistant over their own reports and a place to train models, without creating a single network path out of the building.
What could not be compromised
- No network connection of any kind from the high side to the outside, including for updates, licensing or telemetry.
- Every file that enters must be traceable to a named requester, a reviewer and a scan result.
- Users only see answers drawn from reports they are already cleared to read.
- Software must run fully offline: no licence servers that phone home, no cloud model APIs.
- Security incidents must still be reportable within CERT-In’s 6-hour window, with evidence from the audit store.
Four ways to give researchers AI, compared
Each option was scored against the security accreditation requirements first, and only then on cost and usefulness.
Everything crosses the gap one way, and is checked twice
The platform sits entirely on the high side. Models, packages and patches are fetched on a separate low-side network, scanned, bundled and signed, then pushed across a data diode whose hardware can only transmit in one direction. On arrival they wait in quarantine for a second scan before anything can use them. Nothing on the high side has a route back out.
| Building block | Why it is there |
|---|---|
| 1 Low-side staging and review | A small, separate network where approved staff download models, Python and OS packages and drivers. Each item is hashed against the publisher’s checksum, scanned, checked for licence terms and export conditions, and tied to a named request. |
| 2 Bundle signing | Approved items are packed into a bundle with a manifest, a software bill of materials and hashes, then signed with a key held in a hardware module. The high side rejects any bundle whose signature or hashes do not match. |
| 3 Data diode and media transfer | A hardware diode with a single fibre that can only transmit, so a return path is physically impossible. Write-once media is the fallback for very large or urgent items, under a two-person rule with a logged chain of custody. |
| 4 Quarantine | Every bundle lands in quarantine and is scanned again with different engines from the low side. It is held for 72 hours so newly published threat signatures can catch up, then released by a second reviewer. |
| 5 Model registry and update mirror | Signed, versioned models with their evaluation results, plus an internal mirror for Python, container images, OS packages and GPU drivers. Nothing on the high side installs from anywhere else. |
| 6 Report assistant and search index | Retrieval over the report library: documents are split into passages, embedded and stored with their clearance labels. The assistant answers only from passages the user may read, and cites the report and page for every claim. |
| 7 GPU cluster | Four nodes of eight GPUs on a 400G InfiniBand fabric with parallel NVMe storage. Two nodes serve models, two run fine-tuning jobs through a Slurm queue, and the split can shift as demand changes. |
| 8 Identity and audit | Smart-card sign-in with need-to-know groups from the existing directory. Every prompt, answer, document retrieved, job and import is written to write-once storage and fed to the high-side SIEM. |
Sized from the report library and how researchers actually work
A two-week survey and a pilot index of 5,000 reports gave the numbers below. Usage is assumed to grow to about 6,000 assistant questions on a busy day once the whole laboratory is onboarded.
| Item | Figure | Basis |
|---|---|---|
| Report library | ~1.2 lakh documents, ~4 TB | Average 40 pages, mostly scanned PDFs needing OCR |
| Search index | ~1.4 crore passages, ~60 GB of embeddings | 1.2 lakh x 40 pages x 3 passages, 1,024 dimensions |
| Peak assistant load | ~6,000 questions a day, ~40 at once | 350 users, about 17 questions each on a heavy day |
| Serving GPUs | 2 nodes, 16 GPUs | 70B-class model at 8-bit on 2 GPUs per replica, 6 replicas, plus embedding and re-ranking |
| Fine-tuning GPUs | 2 nodes, 16 GPUs | LoRA on 70B-class or full tuning up to 8B; queue measured at 3 to 4 jobs a week |
| Import volume | ~2 TB a month through the diode | A 70B model is ~140 GB; at ~1 Gbps effective it crosses in about 20 minutes |
| Parallel storage | ~500 TB usable NVMe | Index, datasets, checkpoints and 2 years of model versions |
Budget for hardware, diode, scanning and software is in the range of ₹22 to 28 crore, of which the GPU nodes are about two thirds. Final GPU counts are confirmed by benchmarking the chosen models on the pilot index before ordering.
Five stages, each signed off by the security accreditation board
On an isolated network the order matters: the import path is built and proven before the platform, because every later component arrives through it.
Design and accreditation plan
Weeks 1 to 6
Threat model, data classification mapping, import procedure and audit design agreed with the laboratory’s security board. Hardware ordered.
Gate: Accreditation board approves the design and the import procedure.
Import path
Weeks 5 to 12
Low-side staging, signing, data diode, media station and quarantine built. Test bundles, including deliberately bad ones, are pushed across.
Gate: Every bad bundle rejected, every good one traceable end to end.
High-side platform
Weeks 10 to 20
GPU cluster, storage, registry, update mirror, identity and audit built from imported, signed artefacts only.
Gate: Platform rebuilt from scratch using only the mirror, with no manual installs.
Pilot with two research groups
Weeks 21 to 28
About 40 users. Report library indexed with clearance labels. Answer quality and access controls tested on real questions.
Gate: Zero clearance leaks in red-team testing; 85% of answers rated useful.
Laboratory-wide rollout
Weeks 29 to 34
Groups onboarded in batches with training. Fine-tuning queue opened. Monthly import calendar set.
Gate: Accreditation renewed with the platform in scope.
What could go wrong on an isolated network, and what is already in the plan
| Risk | What could happen | How the design handles it |
|---|---|---|
| Malicious file crosses the gap | A poisoned package or model carries hidden code | Two independent scans on different engines, signature checks, a 72-hour hold, and models loaded only in safe formats with no executable code. |
| Clearance leak through the assistant | An answer draws on a report the user cannot read | Clearance labels filter retrieval before the model sees anything. Red-team tests at every release, and every retrieval is logged. |
| Stale software | Patches lag because imports are slow | A fixed monthly import calendar, plus an emergency path with two-person approval for critical fixes. |
| Data diode failure | Imports stop for weeks | Write-once media path runs on the same signing and quarantine process, so it can take over the same day. |
| Skills on a closed network | Engineers cannot search the web for fixes | An offline documentation mirror, runbooks written during build, and a support contract with on-site response. |
What was optimised
Paths out of the building
The diode is one-way in hardware, not in configuration. There is nothing to misconfigure.
Independent scanning
Different engines on each side of the gap, so one missed signature does not let a file through.
Large-model import
A 140 GB model crosses the diode in about 20 minutes instead of days of media handling.
Shared GPU split
Serving and fine-tuning share one cluster, and nodes move between them as demand shifts.
Answers with citations
Every answer names the report and page, so researchers check the source, not the model.
One place to install from
All software on the high side comes from one signed mirror, so a rebuild is repeatable.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Time to find prior work | Hours to days | Minutes, with citations to the source report |
| Repeated tests from lost results | Several a year | Down sharply as earlier results become findable |
| Model and package import | Not possible in a controlled way | Monthly calendar, 1 to 3 days from request to use |
| Network paths to the outside | None | Still none, proven by design and testing |
| Audit evidence | File-share logs only | Every import, prompt, retrieval and job recorded on write-once storage |
| Fine-tuning for research groups | Not available | 3 to 4 jobs a week on the shared queue |
Targets are measured in the pilot on the laboratory’s own reports. Answer quality depends heavily on OCR quality of older scanned documents, which is assessed during the pilot indexing.
What a team needs to deliver this
Air-gapped architecture
One-way transfer design, import procedures, signing and chain of custody that stand up to accreditation.
GPU cluster design
Node, fabric and storage sizing, Slurm and Kubernetes scheduling, and offline driver management.
LLM serving and fine-tuning
Model selection, quantisation, replica sizing and LoRA workflows from measured workloads.
Retrieval with access control
Passage-level clearance labels, filtered retrieval and red-team testing for leaks.
Software supply chain security
SBOMs, signature checks, multi-engine scanning and safe model formats.
Security operations
Write-once audit, SIEM integration and incident reporting on a closed network.
Other scenarios
Shared GPU cluster for research
Nine universities share one 128-GPU cluster, with fair-share scheduling, group quotas and clear usage reporting for every member.
Read →Cloud repatriationPublic cloud to private cloud
A B2B SaaS company moves its steady workloads off public cloud to a private cloud, and keeps burst capacity where it is cheap.
Read →Hybrid cloudHybrid cloud for omnichannel retail
An omnichannel retailer keeps ERP and stores on-premises and bursts its online storefront to public cloud for 10x sale-day traffic.
Read →Need AI on a network that can never touch the internet?
Tell us about your isolated environment, the data you hold and what your accreditation requires. We will come back with a plain view of how models and updates could cross safely, what the platform needs, and how to prove nothing goes back out.