Home / Deployment scenarios / Shared GPU cluster for research
Reference deployment Research computing

One GPU cluster, shared fairly by nine universities

Nine universities in a research consortium each wanted GPUs for AI, climate, genomics and materials science. Bought separately, they would have nine small clusters, nine sets of staff and a lot of idle hardware. This design pools the money into one 128-GPU cluster at a host campus, and puts most of the effort where shared clusters usually fail: who gets what, when, and how everyone can see it is fair.

SectorHigher education and research
Cluster16 nodes, 128 GPUs
Users~60 groups, ~900 people
ModelShared, cost recovery
The situation

Plenty of demand, scattered hardware, and long queues for the wrong people

The consortium brings together nine public and private universities. Between them, about 60 research groups need GPUs: deep learning, climate and weather modelling, protein structure work, computational chemistry and a growing number of language-model projects. Today they use a mix of departmental workstations, a few ageing four-GPU servers and cloud credits that run out mid-project.

A survey of the groups found that the GPUs they already own sit idle more than half the time, while other groups wait weeks for access or cut their experiments down to fit. PhD students lose months. Grant reviewers have started asking why the same equipment appears in several proposals.

A joint funding round makes about ₹85 crore available over five years for a shared facility, on one condition: every member must be able to see that it gets a share in line with what it puts in, and the facility must recover its running costs from usage.

What could not be compromised

  • Fair access that each university can verify: shares proportional to funding, visible in monthly reports.
  • Batch jobs for large training and simulation runs, and interactive notebooks for teaching and exploration, on one facility.
  • Research data stays on infrastructure the consortium controls; some health and genomics datasets fall under the DPDP Act and ethics approvals.
  • The host data centre has limited air-cooling capacity: around 20 kW per rack today.
  • A small central team of five engineers, plus research software engineers at member campuses.
Options weighed

Four ways to give researchers GPUs, compared on five-year cost and fairness

Each option was costed on the same demand forecast: about 9 lakh GPU-hours a year, rising as new groups join.

OptionWhat worksWhat does notVerdict
Each university buys its ownLocal control, simple governance.Nine small clusters, nine sets of staff, low utilisation. No member can afford large multi-node jobs.Rejected
Commercial GPU cloud creditsNo hardware, no facility work.At sustained use the five-year cost is two to three times higher. Budgets run out mid-project and large datasets are slow and costly to move.Rejected
Shared cluster, Slurm batch onlyProven model for research computing, simple to run.Notebook users and teaching sit in the batch queue or lock GPUs for hours.Considered
Shared cluster, Slurm batch plus a small Kubernetes partition for notebooks and servicesLarge jobs and interactive work each get a fitting scheduler, on one fabric and one storage system.Two schedulers to run, so accounting has to be joined up from day one.Chosen
Target architecture

Fair-share scheduling in the middle, everything else built around it

Fourteen GPU nodes run batch work under Slurm with fair-share. Two nodes sit in a Kubernetes partition with their GPUs sliced into smaller pieces for notebooks, teaching and small services. Both draw on the same tiered storage over the same InfiniBand fabric, and both report usage into one accounting database that drives the monthly reports and chargeback.

Scroll sideways to see the whole diagram →
MEMBER UNIVERSITIES AND RESEARCH GROUPSSHARED RESEARCH CLUSTER, HOST UNIVERSITY DATA CENTREGOVERNANCE AND BILLINGExternal datasetspublic archives, partner labsResearch groups~60 groups across 9 universities, ~900 active usersMember universitiesfunding share, finance officeData transfer nodes2 x 100G to research networkNotebook portalJupyter, RStudio, desktopsLogin nodesSSH, compile, submitProject storage~3 PB, quota per groupKubernetes2 nodes, sliced GPUsSlurm schedulerfair share, QOS, preemptionParallel scratch~1 PB NVMe, 30-day purgeGPU nodes: 16 x 8 GPUs128 GPUs, 400G InfiniBand, 14 batch + 2 interactiveArchive tiertape-backed, 7 yearsMonitoringGPU, power, coolant tempsDirect liquid cooling4 racks, ~42 kW eachAllocation committeequarterly shares and quotasSlurm accountingGPU-hours per groupUsage reportsper group and per universityChargebackcost recovery per memberbulk copystage in1notebooksSSHsessions2batch jobsfunding shareinteractiveschedulestelemetrycoolantnightly6monthly354Data / replicationScheduled copyUser or API trafficControl / API callLogging / management
Numbered flows: (1) researchers open notebooks through the web portal, (2) batch jobs are submitted from the login nodes to Slurm, (3) the allocation committee sets each university’s share and each group’s quota, (4) jobs read and write the parallel scratch tier at full fabric speed, (5) every job record lands in the accounting database, (6) usage reports become a monthly chargeback per member.
Building blockWhy it is there
1 GPU nodesSixteen nodes with eight GPUs each, joined by 400G InfiniBand in a non-blocking fat tree so a job can span several nodes at full speed. Fourteen run batch work, two serve the interactive partition.
2 Slurm with fair-shareEach university gets a share in proportion to its funding, split further between its groups. Recent heavy users drop in priority, light users rise. Quality-of-service levels cap job length and size, and short jobs can backfill around large ones.
3 Kubernetes partitionTwo nodes with GPUs split into smaller slices for notebooks, coursework and small inference services. Sessions time out when idle, so GPUs are not held overnight by an open browser tab.
4 Notebook portal and login nodesA web portal for Jupyter, RStudio and remote desktops, and login nodes for SSH users who submit scripts. Sign-in goes through each university’s own identity system.
5 Tiered storageA 1 PB NVMe parallel scratch tier for active jobs, purged after 30 days; about 3 PB of project storage with quotas per group; and a tape-backed archive for finished projects, kept seven years.
6 Data transfer nodesDedicated servers on the national research network for moving large datasets in and out, so bulk copies never slow the login nodes or the compute fabric.
7 Accounting and chargebackOne database collects GPU-hours, storage use and job records from both schedulers. Monthly reports show each group and university its usage against its share, and drive cost recovery.
8 Direct liquid coolingCold plates on GPUs and CPUs carry most of the heat to coolant distribution units. Four racks at about 42 kW each fit a room that could only take about 20 kW per rack by air.
Sizing, worked out

Sized from the groups’ own forecasts, then checked against power

Each group was asked for GPU-hours, largest job size and data volume for the next two years. Their totals were discounted by a third, which matches how research forecasts usually compare with real use.

ItemFigureBasis
GPU demand~9 lakh GPU-hours a year~13.5 lakh forecast by groups, less one third
GPU count128 GPUs (16 nodes x 8)128 x 8,760 hours = ~11.2 lakh hours; 9 lakh is ~80% utilisation
Largest routine job32 GPUs for up to 72 hoursFour nodes; fat-tree fabric keeps multi-node jobs at full speed
Scratch tier~1 PB NVMe, ~150 GB/sLargest training sets plus checkpoints; ~1.2 GB/s per GPU at full load
Project storage~3 PB usable60 groups at an average 40 TB, plus 20% headroom
Power and heat~200 kW IT load16 nodes x ~10.2 kW, plus storage, fabric and services; 4 racks at ~42 kW
Running cost to recover~₹6 to 7 crore a yearPower at PUE 1.3, support contracts and staff; hardware is grant-funded

At about 9 lakh GPU-hours, cost recovery works out to roughly ₹65 to 80 per GPU-hour, well below commercial rates because hardware is paid from the joint grant and only running costs are recovered. Hardware, fabric, storage and cooling come to about ₹70 to 78 crore of the five-year budget.

How it is delivered

From an empty room to a full queue in five stages

Facility work comes first, because liquid cooling and power upgrades set the critical path, not servers.

1

Governance and design

Months 1 to 2

Consortium agreement on shares, quotas, priorities and the chargeback rate. Facility survey and cooling design. Hardware tender issued.

Gate: Every member signs the allocation policy and the rate card.

2

Facility readiness

Months 2 to 5

Power upgrade, coolant loop, CDUs and rack positions installed and pressure-tested.

Gate: Cooling loop runs at full simulated heat load for 72 hours.

3

Cluster build

Months 4 to 6

Nodes, fabric, storage, Slurm, Kubernetes, portal and accounting built from code. Burn-in on every GPU.

Gate: Benchmarks within 5% of expected; failure tests passed.

4

Early users

Months 7 to 8

Ten groups from five universities run real work. Fair-share weights and quotas tuned on real queues.

Gate: Median queue wait under 4 hours for standard jobs.

5

Full service

Month 9 onwards

All groups onboarded with training. First monthly reports and chargeback issued. Quarterly allocation reviews begin.

Gate: Every member accepts its first usage report.

Way back: Scheduler and storage settings are kept in version control, so a bad change to shares, limits or partitions is reversed in minutes. During early use, groups keep their existing local servers until their workloads have run successfully on the cluster.
Risks, handled up front

What usually goes wrong with shared clusters, and how the design handles it

RiskWhat could happenHow the design handles it
Perceived unfairnessA member feels a few large groups take everythingShares fixed by agreement, fair-share decay, published monthly reports, and a committee that reviews disputes quarterly.
GPUs held idleInteractive sessions or stuck jobs hold GPUs without using themIdle notebooks time out; jobs with near-zero GPU use for an hour are flagged to the owner and then ended.
Storage fills upScratch or project storage runs out mid-semesterHard quotas per group, 30-day scratch purge, and warnings at 80% sent to the group lead.
Cooling incidentA coolant leak or CDU failureN+1 CDUs, leak detection under each rack, and automatic job draining and shutdown before temperatures reach limits.
Sensitive research dataHealth or genomics data reachable by other groupsRestricted projects get separate storage, access tied to ethics approvals, and no notebook access to their data from shared partitions.
What was optimised

What was optimised

~80%

Utilisation target

One pool replaces nine half-idle clusters, with backfill filling the gaps around large jobs.

1 rate

Simple chargeback

A single GPU-hour and TB-month rate, published in advance, so groups can budget in grant proposals.

7 slices

Smaller GPUs for teaching

Sliced GPUs give a class of students their own piece of a GPU without tying up whole cards.

42 kW

Rack density

Liquid cooling fits the cluster in four racks instead of eight or more air-cooled ones.

3 tiers

Storage by purpose

Fast scratch only for active jobs, larger project storage, and low-cost archive for finished work.

1 database

One view of usage

Batch and interactive use land in the same accounting records, so reports add up.

Outcomes

What the design is built to deliver

MeasureBeforeDesign target
GPU utilisationUnder 50% on scattered servers75 to 85% across the shared cluster
Median wait for a standard jobDays to weeks, or no accessUnder 4 hours
Largest practical job4 GPUs on one server32 GPUs routinely, more by arrangement
Cost per GPU-hour to groupsCloud rates, or hidden in department budgets~₹65 to 80, running costs only
Visibility of fair useNoneMonthly report per group and per university
StaffPart-time admins at each campusOne central team of five, with campus research software engineers

Targets are reviewed with the allocation committee after the early-user stage. Wait times depend heavily on how quickly demand grows and on the mix of large and small jobs.

Skills this draws on

What a team needs to deliver this

HPC and GPU cluster design

Node, fabric and storage sizing for mixed training and simulation work.

Slurm and fair-share policy

Shares, quality of service, backfill and preemption that match a consortium agreement.

Kubernetes for research

GPU slicing, notebook platforms and idle-session controls.

Parallel and tiered storage

Scratch, project and archive tiers with quotas and purge policies.

Data centre power and liquid cooling

Direct liquid cooling design, CDU redundancy and safe shutdown under fault.

Accounting and cost recovery

Usage data from several schedulers joined into reports and a rate card members trust.

About this page. This is a reference deployment: a worked design built from requirements we see repeatedly in this kind of organisation. It is not a description of a specific client. Figures are design targets and planning estimates; real numbers depend on your workloads and are confirmed during assessment. We are glad to walk through how it would apply to your environment.

Planning a shared GPU facility across departments or institutions?

Tell us who will share it, roughly how much demand you expect and what your data centre can take. We will come back with a plain first view of sizing, cooling, scheduling policy and how usage could be shared and recovered fairly.