One GPU cluster, shared fairly by nine universities
Nine universities in a research consortium each wanted GPUs for AI, climate, genomics and materials science. Bought separately, they would have nine small clusters, nine sets of staff and a lot of idle hardware. This design pools the money into one 128-GPU cluster at a host campus, and puts most of the effort where shared clusters usually fail: who gets what, when, and how everyone can see it is fair.
Plenty of demand, scattered hardware, and long queues for the wrong people
The consortium brings together nine public and private universities. Between them, about 60 research groups need GPUs: deep learning, climate and weather modelling, protein structure work, computational chemistry and a growing number of language-model projects. Today they use a mix of departmental workstations, a few ageing four-GPU servers and cloud credits that run out mid-project.
A survey of the groups found that the GPUs they already own sit idle more than half the time, while other groups wait weeks for access or cut their experiments down to fit. PhD students lose months. Grant reviewers have started asking why the same equipment appears in several proposals.
A joint funding round makes about ₹85 crore available over five years for a shared facility, on one condition: every member must be able to see that it gets a share in line with what it puts in, and the facility must recover its running costs from usage.
What could not be compromised
- Fair access that each university can verify: shares proportional to funding, visible in monthly reports.
- Batch jobs for large training and simulation runs, and interactive notebooks for teaching and exploration, on one facility.
- Research data stays on infrastructure the consortium controls; some health and genomics datasets fall under the DPDP Act and ethics approvals.
- The host data centre has limited air-cooling capacity: around 20 kW per rack today.
- A small central team of five engineers, plus research software engineers at member campuses.
Four ways to give researchers GPUs, compared on five-year cost and fairness
Each option was costed on the same demand forecast: about 9 lakh GPU-hours a year, rising as new groups join.
Fair-share scheduling in the middle, everything else built around it
Fourteen GPU nodes run batch work under Slurm with fair-share. Two nodes sit in a Kubernetes partition with their GPUs sliced into smaller pieces for notebooks, teaching and small services. Both draw on the same tiered storage over the same InfiniBand fabric, and both report usage into one accounting database that drives the monthly reports and chargeback.
| Building block | Why it is there |
|---|---|
| 1 GPU nodes | Sixteen nodes with eight GPUs each, joined by 400G InfiniBand in a non-blocking fat tree so a job can span several nodes at full speed. Fourteen run batch work, two serve the interactive partition. |
| 2 Slurm with fair-share | Each university gets a share in proportion to its funding, split further between its groups. Recent heavy users drop in priority, light users rise. Quality-of-service levels cap job length and size, and short jobs can backfill around large ones. |
| 3 Kubernetes partition | Two nodes with GPUs split into smaller slices for notebooks, coursework and small inference services. Sessions time out when idle, so GPUs are not held overnight by an open browser tab. |
| 4 Notebook portal and login nodes | A web portal for Jupyter, RStudio and remote desktops, and login nodes for SSH users who submit scripts. Sign-in goes through each university’s own identity system. |
| 5 Tiered storage | A 1 PB NVMe parallel scratch tier for active jobs, purged after 30 days; about 3 PB of project storage with quotas per group; and a tape-backed archive for finished projects, kept seven years. |
| 6 Data transfer nodes | Dedicated servers on the national research network for moving large datasets in and out, so bulk copies never slow the login nodes or the compute fabric. |
| 7 Accounting and chargeback | One database collects GPU-hours, storage use and job records from both schedulers. Monthly reports show each group and university its usage against its share, and drive cost recovery. |
| 8 Direct liquid cooling | Cold plates on GPUs and CPUs carry most of the heat to coolant distribution units. Four racks at about 42 kW each fit a room that could only take about 20 kW per rack by air. |
Sized from the groups’ own forecasts, then checked against power
Each group was asked for GPU-hours, largest job size and data volume for the next two years. Their totals were discounted by a third, which matches how research forecasts usually compare with real use.
| Item | Figure | Basis |
|---|---|---|
| GPU demand | ~9 lakh GPU-hours a year | ~13.5 lakh forecast by groups, less one third |
| GPU count | 128 GPUs (16 nodes x 8) | 128 x 8,760 hours = ~11.2 lakh hours; 9 lakh is ~80% utilisation |
| Largest routine job | 32 GPUs for up to 72 hours | Four nodes; fat-tree fabric keeps multi-node jobs at full speed |
| Scratch tier | ~1 PB NVMe, ~150 GB/s | Largest training sets plus checkpoints; ~1.2 GB/s per GPU at full load |
| Project storage | ~3 PB usable | 60 groups at an average 40 TB, plus 20% headroom |
| Power and heat | ~200 kW IT load | 16 nodes x ~10.2 kW, plus storage, fabric and services; 4 racks at ~42 kW |
| Running cost to recover | ~₹6 to 7 crore a year | Power at PUE 1.3, support contracts and staff; hardware is grant-funded |
At about 9 lakh GPU-hours, cost recovery works out to roughly ₹65 to 80 per GPU-hour, well below commercial rates because hardware is paid from the joint grant and only running costs are recovered. Hardware, fabric, storage and cooling come to about ₹70 to 78 crore of the five-year budget.
From an empty room to a full queue in five stages
Facility work comes first, because liquid cooling and power upgrades set the critical path, not servers.
Governance and design
Months 1 to 2
Consortium agreement on shares, quotas, priorities and the chargeback rate. Facility survey and cooling design. Hardware tender issued.
Gate: Every member signs the allocation policy and the rate card.
Facility readiness
Months 2 to 5
Power upgrade, coolant loop, CDUs and rack positions installed and pressure-tested.
Gate: Cooling loop runs at full simulated heat load for 72 hours.
Cluster build
Months 4 to 6
Nodes, fabric, storage, Slurm, Kubernetes, portal and accounting built from code. Burn-in on every GPU.
Gate: Benchmarks within 5% of expected; failure tests passed.
Early users
Months 7 to 8
Ten groups from five universities run real work. Fair-share weights and quotas tuned on real queues.
Gate: Median queue wait under 4 hours for standard jobs.
Full service
Month 9 onwards
All groups onboarded with training. First monthly reports and chargeback issued. Quarterly allocation reviews begin.
Gate: Every member accepts its first usage report.
What usually goes wrong with shared clusters, and how the design handles it
| Risk | What could happen | How the design handles it |
|---|---|---|
| Perceived unfairness | A member feels a few large groups take everything | Shares fixed by agreement, fair-share decay, published monthly reports, and a committee that reviews disputes quarterly. |
| GPUs held idle | Interactive sessions or stuck jobs hold GPUs without using them | Idle notebooks time out; jobs with near-zero GPU use for an hour are flagged to the owner and then ended. |
| Storage fills up | Scratch or project storage runs out mid-semester | Hard quotas per group, 30-day scratch purge, and warnings at 80% sent to the group lead. |
| Cooling incident | A coolant leak or CDU failure | N+1 CDUs, leak detection under each rack, and automatic job draining and shutdown before temperatures reach limits. |
| Sensitive research data | Health or genomics data reachable by other groups | Restricted projects get separate storage, access tied to ethics approvals, and no notebook access to their data from shared partitions. |
What was optimised
Utilisation target
One pool replaces nine half-idle clusters, with backfill filling the gaps around large jobs.
Simple chargeback
A single GPU-hour and TB-month rate, published in advance, so groups can budget in grant proposals.
Smaller GPUs for teaching
Sliced GPUs give a class of students their own piece of a GPU without tying up whole cards.
Rack density
Liquid cooling fits the cluster in four racks instead of eight or more air-cooled ones.
Storage by purpose
Fast scratch only for active jobs, larger project storage, and low-cost archive for finished work.
One view of usage
Batch and interactive use land in the same accounting records, so reports add up.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| GPU utilisation | Under 50% on scattered servers | 75 to 85% across the shared cluster |
| Median wait for a standard job | Days to weeks, or no access | Under 4 hours |
| Largest practical job | 4 GPUs on one server | 32 GPUs routinely, more by arrangement |
| Cost per GPU-hour to groups | Cloud rates, or hidden in department budgets | ~₹65 to 80, running costs only |
| Visibility of fair use | None | Monthly report per group and per university |
| Staff | Part-time admins at each campus | One central team of five, with campus research software engineers |
Targets are reviewed with the allocation committee after the early-user stage. Wait times depend heavily on how quickly demand grows and on the mix of large and small jobs.
What a team needs to deliver this
HPC and GPU cluster design
Node, fabric and storage sizing for mixed training and simulation work.
Slurm and fair-share policy
Shares, quality of service, backfill and preemption that match a consortium agreement.
Kubernetes for research
GPU slicing, notebook platforms and idle-session controls.
Parallel and tiered storage
Scratch, project and archive tiers with quotas and purge policies.
Data centre power and liquid cooling
Direct liquid cooling design, CDU redundancy and safe shutdown under fault.
Accounting and cost recovery
Usage data from several schedulers joined into reports and a rate card members trust.
Other scenarios
Air-gapped private AI platform
A public research laboratory runs its own AI platform with no internet connection at all, fed only through a one-way data diode.
Read →Cloud repatriationPublic cloud to private cloud
A B2B SaaS company moves its steady workloads off public cloud to a private cloud, and keeps burst capacity where it is cheap.
Read →Hybrid cloudHybrid cloud for omnichannel retail
An omnichannel retailer keeps ERP and stores on-premises and bursts its online storefront to public cloud for 10x sale-day traffic.
Read →Planning a shared GPU facility across departments or institutions?
Tell us who will share it, roughly how much demand you expect and what your data centre can take. We will come back with a plain first view of sizing, cooling, scheduling policy and how usage could be shared and recovered fairly.