GPUs accelerate the matrix calculations at the heart of machine learning, and are often the largest cost in AI projects.
Training Versus Inference
- Training needs large memory and fast interconnects for multi-GPU jobs.
- Inference prioritises cost per request, latency and memory to hold models.
Sizing
GPU memory must hold model weights, activations and, for language models, the key-value cache. Quantisation reduces requirements.
Cloud Versus On-Premises
- Cloud: flexible, no upfront investment, but can be expensive at sustained high usage and capacity may be scarce.
- On-premises: lower long-run cost at high utilisation, but requires capital, expertise and power and cooling.
Improving Utilisation
- Batch requests.
- Share GPUs across smaller workloads.
- Schedule jobs to avoid idle time.
- Monitor utilisation; idle GPUs are expensive.
Alternatives
Managed model APIs avoid GPU management entirely. Other accelerators exist too.
Cost Control
Use spot or preemptible instances for fault-tolerant training, with checkpointing.