Storage for AI
GPUs can only work as fast as data reaches them. Storage for AI has two jobs: feed training data quickly, and absorb very large checkpoints without stopping the GPUs for long. General-purpose file servers are rarely up to either.
Three tiers
| Tier | Holds | Typical technology |
|---|---|---|
| Local NVMe in each server | Scratch space and cached data | NVMe drives inside the GPU servers |
| Parallel file system | Active datasets and checkpoints | Lustre, IBM Storage Scale (GPFS), WEKA, VAST, DDN |
| Object storage | Raw data, archives, finished models | S3-compatible storage such as Ceph RGW or MinIO |
Checkpoint maths
During training, the full state of the model is saved regularly so a failure does not lose days of work. With mixed-precision training and the Adam optimizer, that state is roughly 16 bytes per parameter.
| Model | Checkpoint size (approx.) | Time to write at 10 GB/s | At 50 GB/s |
|---|---|---|---|
| 8B | ~130 GB | ~13 seconds | ~3 seconds |
| 70B | ~1.1 TB | ~2 minutes | ~22 seconds |
| 405B | ~6.5 TB | ~11 minutes | ~2 minutes |
Multiply the write time by how often you checkpoint, and you have GPU time spent waiting. That number, more than capacity, is what decides how fast the storage needs to be.
Getting it right
- Size for throughput first, then capacity. A small, fast tier in front of large, cheap object storage is usually better value than a large fast tier.
- Use asynchronous checkpointing where your framework supports it. The GPUs carry on while the checkpoint is written in the background.
- Keep many small files out of the fast tier. Pack datasets into larger shards. Millions of tiny files slow any file system down.
- Back up what matters. Datasets and final models need backup. Intermediate checkpoints usually do not.
More AI infrastructure guides: Training cluster · Inference platform · GPU network fabric · Power and cooling · FAQ · Use case: private LLM platform
Planning a GPU cluster?
Tell us the models you want to run or train, how many users, and where it will be hosted. We will come back with a first sizing: GPUs, servers, network, storage, and the power and cooling your data centre will need.