Home / Deployment scenarios / Air-gapped private AI platform
Reference deployment Private AI

A private AI platform with no connection to the outside world

A public research laboratory holds decades of technical reports that fall under export control. Its scientists want the same AI help their peers elsewhere take for granted, but the data cannot sit on any network that touches the internet. This design builds a complete AI platform inside the laboratory’s isolated network, and treats every model, package and patch that crosses the air gap as something to be checked, signed and recorded.

SectorPublic research
Platform32 GPUs, 4 nodes
Programme8 months to lab-wide use
ModelFully air-gapped
The situation

Valuable knowledge, locked in reports nobody can search

The laboratory has about 350 scientists and engineers cleared to work on its classified network. Over thirty years they have written roughly 1.2 lakh technical reports, test records and design notes. Much of it is export-controlled technical data under the SCOMET list, and all of it lives on a network with no physical connection to the internet.

Finding prior work is slow. Search on the isolated network is a basic keyword tool over file shares, and senior staff estimate that engineers spend a day or more a week looking for earlier results, or repeating tests because they could not find them. Several research groups also want to fine-tune models on their own data, which today means writing code on paper notes and carrying it in.

Public AI services are out of the question, and so is a normal on-premises platform that pulls updates from the internet. The director’s brief is direct: give researchers a capable assistant over their own reports and a place to train models, without creating a single network path out of the building.

What could not be compromised

  • No network connection of any kind from the high side to the outside, including for updates, licensing or telemetry.
  • Every file that enters must be traceable to a named requester, a reviewer and a scan result.
  • Users only see answers drawn from reports they are already cleared to read.
  • Software must run fully offline: no licence servers that phone home, no cloud model APIs.
  • Security incidents must still be reportable within CERT-In’s 6-hour window, with evidence from the audit store.
Options weighed

Four ways to give researchers AI, compared

Each option was scored against the security accreditation requirements first, and only then on cost and usefulness.

OptionWhat worksWhat does notVerdict
Public or government cloud AI with private endpointsFast to start, no hardware.Export-controlled data would leave the laboratory. Fails accreditation outright.Rejected
On-premises platform with filtered internet for updatesEasy model and patch updates.A filtered path is still a path. The isolated network would lose its status.Rejected
Air-gapped platform, imports on removable media onlyNo network path at all.Slow for 100+ GB models, many manual steps, hard to audit consistently.Kept as fallback
Air-gapped platform, imports through a one-way data diode plus mediaPhysically one-way. Every import scanned, signed and logged. Large models arrive in hours.Diode and two-stage scanning add cost and process. Imports need planning.Chosen
Target architecture

Everything crosses the gap one way, and is checked twice

The platform sits entirely on the high side. Models, packages and patches are fetched on a separate low-side network, scanned, bundled and signed, then pushed across a data diode whose hardware can only transmit in one direction. On arrival they wait in quarantine for a second scan before anything can use them. Nothing on the high side has a route back out.

Scroll sideways to see the whole diagram →
LOW SIDE: CONNECTEDAIR GAP: ONE WAY INHIGH SIDE: PEOPLE AND CONTROLSHIGH SIDE: PRIVATE AI PLATFORM, NO INTERNETUpstream sourcesmodel hubs, OS vendorsLow-side stagingdownload, hash, pre-scanBundle and signmanifest, SBOM, hashesImport reviewlicence, export checkData diodeone-way fibre onlyMedia transfertwo-person, loggedNo return pathno link back outAudit and SIEMWORM logs, 1 yr onlineIdentity and MFAsmart card, clearancesResearcherscleared staff, ~350Report libraryinternal reports, ~4 TBSearch indexembeddings + clearancesReport assistantanswers with citationsQuarantinesecond scan, 72 h holdModel registrysigned model versionsInference serving70B-class, 16 GPUsUpdate mirrorOS, drivers, packagesFine-tuning jobsSlurm queue, 16 GPUsGPU cluster: 4 nodes x 8 GPUs32 GPUs, 400G InfiniBand, parallel NVMe storagepatchesnew versions4questionspromptsruns onjobs12356Data / replicationControl / API callScheduled copyUser or API trafficLogging / management
Numbered flows: (1) a signed bundle crosses the one-way data diode, with write-once media as the fallback path, (2) after a second scan and a 72-hour hold it is released to the model registry, (3) approved models are deployed to inference serving, (4) cleared researchers ask questions through the report assistant, (5) the assistant retrieves only passages the user is cleared to see, (6) OS and driver patches reach the GPU cluster from the offline update mirror. The high side has no outbound path of any kind.
Building blockWhy it is there
1 Low-side staging and reviewA small, separate network where approved staff download models, Python and OS packages and drivers. Each item is hashed against the publisher’s checksum, scanned, checked for licence terms and export conditions, and tied to a named request.
2 Bundle signingApproved items are packed into a bundle with a manifest, a software bill of materials and hashes, then signed with a key held in a hardware module. The high side rejects any bundle whose signature or hashes do not match.
3 Data diode and media transferA hardware diode with a single fibre that can only transmit, so a return path is physically impossible. Write-once media is the fallback for very large or urgent items, under a two-person rule with a logged chain of custody.
4 QuarantineEvery bundle lands in quarantine and is scanned again with different engines from the low side. It is held for 72 hours so newly published threat signatures can catch up, then released by a second reviewer.
5 Model registry and update mirrorSigned, versioned models with their evaluation results, plus an internal mirror for Python, container images, OS packages and GPU drivers. Nothing on the high side installs from anywhere else.
6 Report assistant and search indexRetrieval over the report library: documents are split into passages, embedded and stored with their clearance labels. The assistant answers only from passages the user may read, and cites the report and page for every claim.
7 GPU clusterFour nodes of eight GPUs on a 400G InfiniBand fabric with parallel NVMe storage. Two nodes serve models, two run fine-tuning jobs through a Slurm queue, and the split can shift as demand changes.
8 Identity and auditSmart-card sign-in with need-to-know groups from the existing directory. Every prompt, answer, document retrieved, job and import is written to write-once storage and fed to the high-side SIEM.
Sizing, worked out

Sized from the report library and how researchers actually work

A two-week survey and a pilot index of 5,000 reports gave the numbers below. Usage is assumed to grow to about 6,000 assistant questions on a busy day once the whole laboratory is onboarded.

ItemFigureBasis
Report library~1.2 lakh documents, ~4 TBAverage 40 pages, mostly scanned PDFs needing OCR
Search index~1.4 crore passages, ~60 GB of embeddings1.2 lakh x 40 pages x 3 passages, 1,024 dimensions
Peak assistant load~6,000 questions a day, ~40 at once350 users, about 17 questions each on a heavy day
Serving GPUs2 nodes, 16 GPUs70B-class model at 8-bit on 2 GPUs per replica, 6 replicas, plus embedding and re-ranking
Fine-tuning GPUs2 nodes, 16 GPUsLoRA on 70B-class or full tuning up to 8B; queue measured at 3 to 4 jobs a week
Import volume~2 TB a month through the diodeA 70B model is ~140 GB; at ~1 Gbps effective it crosses in about 20 minutes
Parallel storage~500 TB usable NVMeIndex, datasets, checkpoints and 2 years of model versions

Budget for hardware, diode, scanning and software is in the range of ₹22 to 28 crore, of which the GPU nodes are about two thirds. Final GPU counts are confirmed by benchmarking the chosen models on the pilot index before ordering.

How it is delivered

Five stages, each signed off by the security accreditation board

On an isolated network the order matters: the import path is built and proven before the platform, because every later component arrives through it.

1

Design and accreditation plan

Weeks 1 to 6

Threat model, data classification mapping, import procedure and audit design agreed with the laboratory’s security board. Hardware ordered.

Gate: Accreditation board approves the design and the import procedure.

2

Import path

Weeks 5 to 12

Low-side staging, signing, data diode, media station and quarantine built. Test bundles, including deliberately bad ones, are pushed across.

Gate: Every bad bundle rejected, every good one traceable end to end.

3

High-side platform

Weeks 10 to 20

GPU cluster, storage, registry, update mirror, identity and audit built from imported, signed artefacts only.

Gate: Platform rebuilt from scratch using only the mirror, with no manual installs.

4

Pilot with two research groups

Weeks 21 to 28

About 40 users. Report library indexed with clearance labels. Answer quality and access controls tested on real questions.

Gate: Zero clearance leaks in red-team testing; 85% of answers rated useful.

5

Laboratory-wide rollout

Weeks 29 to 34

Groups onboarded in batches with training. Fine-tuning queue opened. Monthly import calendar set.

Gate: Accreditation renewed with the platform in scope.

Way back: Every model and package version stays in the registry and mirror. A bad model or update is rolled back by redeploying the previous signed version, which takes minutes and needs no import. Researchers keep the existing keyword search throughout.
Risks, handled up front

What could go wrong on an isolated network, and what is already in the plan

RiskWhat could happenHow the design handles it
Malicious file crosses the gapA poisoned package or model carries hidden codeTwo independent scans on different engines, signature checks, a 72-hour hold, and models loaded only in safe formats with no executable code.
Clearance leak through the assistantAn answer draws on a report the user cannot readClearance labels filter retrieval before the model sees anything. Red-team tests at every release, and every retrieval is logged.
Stale softwarePatches lag because imports are slowA fixed monthly import calendar, plus an emergency path with two-person approval for critical fixes.
Data diode failureImports stop for weeksWrite-once media path runs on the same signing and quarantine process, so it can take over the same day.
Skills on a closed networkEngineers cannot search the web for fixesAn offline documentation mirror, runbooks written during build, and a support contract with on-site response.
What was optimised

What was optimised

0

Paths out of the building

The diode is one-way in hardware, not in configuration. There is nothing to misconfigure.

2x

Independent scanning

Different engines on each side of the gap, so one missed signature does not let a file through.

20 min

Large-model import

A 140 GB model crosses the diode in about 20 minutes instead of days of media handling.

16 + 16

Shared GPU split

Serving and fine-tuning share one cluster, and nodes move between them as demand shifts.

100%

Answers with citations

Every answer names the report and page, so researchers check the source, not the model.

1 mirror

One place to install from

All software on the high side comes from one signed mirror, so a rebuild is repeatable.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
Time to find prior workHours to daysMinutes, with citations to the source report
Repeated tests from lost resultsSeveral a yearDown sharply as earlier results become findable
Model and package importNot possible in a controlled wayMonthly calendar, 1 to 3 days from request to use
Network paths to the outsideNoneStill none, proven by design and testing
Audit evidenceFile-share logs onlyEvery import, prompt, retrieval and job recorded on write-once storage
Fine-tuning for research groupsNot available3 to 4 jobs a week on the shared queue

Targets are measured in the pilot on the laboratory’s own reports. Answer quality depends heavily on OCR quality of older scanned documents, which is assessed during the pilot indexing.

Skills this draws on

What a team needs to deliver this

Air-gapped architecture

One-way transfer design, import procedures, signing and chain of custody that stand up to accreditation.

GPU cluster design

Node, fabric and storage sizing, Slurm and Kubernetes scheduling, and offline driver management.

LLM serving and fine-tuning

Model selection, quantisation, replica sizing and LoRA workflows from measured workloads.

Retrieval with access control

Passage-level clearance labels, filtered retrieval and red-team testing for leaks.

Software supply chain security

SBOMs, signature checks, multi-engine scanning and safe model formats.

Security operations

Write-once audit, SIEM integration and incident reporting on a closed network.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

Need AI on a network that can never touch the internet?

Tell us about your isolated environment, the data you hold and what your accreditation requires. We will come back with a plain view of how models and updates could cross safely, what the platform needs, and how to prove nothing goes back out.