The short answer
Quick answer: Training a neural network consists almost entirely of one kind of arithmetic: multiplying and adding enormous grids of numbers. Each of those calculations is simple and independent of the others, so they can all be done at once. A CPU has a small number of powerful cores designed to do complicated tasks one after another. A GPU has thousands of simple cores designed to do the same operation on many numbers simultaneously, which is exactly what this workload needs. Modern models have billions of parameters and are trained on trillions of words, so even with GPUs the job takes thousands of chips running for weeks.
What training actually computes
A neural network is layers of numbers called weights. Running data through a layer is a matrix multiplication: a grid of inputs times a grid of weights.
Training repeats three stages, described in how neural networks learn:
- Forward pass: multiply the inputs through every layer to get a prediction.
- Backward pass: multiply back through every layer to work out how each weight should change.
- Update: adjust every weight slightly. See gradient descent explained.
All three are matrix arithmetic. A large language model has billions of weights, and one training step touches every one of them for every piece of input. Training involves an immense number of such steps.
CPU vs GPU
| CPU | GPU | |
|---|---|---|
| Cores | A few to a few dozen, each powerful | Thousands, each simple |
| Designed for | Varied, branching, sequential logic | One operation applied to masses of data |
| Strength | Flexibility, single-thread speed | Throughput on parallel arithmetic |
| Analogy | A handful of expert chefs | A vast kitchen of line cooks all chopping at once |
When you multiply two matrices, every element of the result is its own small calculation, independent of all the others. Nothing needs to wait. A CPU works through them a few at a time. A GPU does thousands at once.
For neural network workloads, that translates into speed-ups of tens to hundreds of times.
Why graphics chips, of all things
A graphics processing unit was built to render video games. Drawing a frame means computing the colour of millions of pixels and transforming millions of triangle corners, using the same matrix maths, all independently, sixty times a second. Chip makers spent decades building hardware that does huge amounts of simple arithmetic in parallel.
Researchers noticed that neural networks needed the same thing. Two developments made it practical:
- General-purpose GPU programming. Nvidia's CUDA platform, released in 2007, let programmers use GPUs for calculations other than graphics.
- A dramatic demonstration. In 2012, a neural network called AlexNet, trained on two consumer gaming graphics cards, won a major image recognition competition by a wide margin. The deep learning boom followed.
Software frameworks such as PyTorch and TensorFlow then made GPU training routine.
Transformers made it scale
Earlier language models read text one word at a time, which limited how much could be done in parallel. The transformer architecture, introduced in Attention Is All You Need, processes all positions of a sequence at once. That fits GPUs perfectly, and it made it worthwhile to train on vastly more data with vastly more hardware. See transformers explained.
Why so many
The models are huge
Billions to hundreds of billions of weights. Every training step reads and updates all of them.
The data is huge
Large language models are trained on trillions of tokens of text. Image and video models consume enormous media collections.
Memory is the hard limit
A GPU has its own memory, called VRAM, and during training it must hold:
- The model's weights.
- The gradients (as many numbers again).
- The optimiser's running statistics (often twice as many again).
- The intermediate values of every layer for the whole batch.
A model with tens of billions of parameters needs far more memory than any single GPU has. It must be spread across many.
Time
Even if one GPU could hold the model, training on it would take lifetimes. Using thousands in parallel brings that down to weeks.
How training is spread across many GPUs
- Data parallelism. Each GPU holds a copy of the model and processes a different slice of the batch. Their gradients are combined before the update.
- Model parallelism. The model itself is split, with different layers or parts on different GPUs.
- Fast interconnects. GPUs must exchange large amounts of data constantly, so clusters use specialised high-speed networking. Communication, not arithmetic, is often the limiting factor.
- Checkpointing. With thousands of chips running for weeks, some will fail. Training regularly saves its state so it can resume.
Adding GPUs does not give a proportional speed-up, because coordination has a cost. Much of the engineering in large-scale training is about keeping every chip busy.
Squeezing more from each chip
- Lower precision. Using 16-bit numbers instead of 32-bit halves the memory needed and speeds up arithmetic, with little effect on results. See floating point maths explained.
- Specialised circuits. Modern AI chips include hardware built specifically for matrix multiplication.
- Software optimisation. Carefully tuned libraries, and techniques that avoid storing or recomputing more than necessary.
It is not only GPUs
GPUs dominate, but they are not the only accelerators:
- Custom chips. Google's TPUs and similar processors from other large technology companies are designed specifically for neural networks.
- On-device chips. Phones and laptops include neural processing units for running small models locally.
- CPUs still matter: they load and prepare data, coordinate the work, and run small models.
Training vs inference
Training creates the model. Inference is using it to answer requests.
| Training | Inference | |
|---|---|---|
| What happens | Forward and backward passes, weight updates | Forward pass only |
| Memory | Weights, gradients, optimiser state, activations | Mostly just weights |
| Scale | Thousands of GPUs for weeks, once per model | From a single chip per request, repeated for every user |
| Can run on smaller hardware | No | Often, especially with compressed models |
Inference is far lighter per request. A model that took a data centre to train may run on one GPU, and small or quantised models run on a laptop or phone. But inference happens constantly, for every user, so across a popular product's lifetime its total computing cost can exceed that of training.
The costs beyond the chips
Large GPU clusters draw enormous amounts of electricity and produce a great deal of heat, so data centres need substantial power supplies and cooling. Demand for AI chips has repeatedly outrun supply. These physical limits, on chips, energy and construction, now shape the pace of AI development as much as algorithms do.
Frequently asked questions
Why are GPUs better than CPUs for AI?
Neural networks need the same simple arithmetic done on huge numbers of values at once. GPUs have thousands of cores built for exactly that; CPUs have a few cores built for sequential, varied work.
Can you train AI without a GPU?
Yes, for small models. For large ones it would take impractically long.
What is VRAM and why does it matter?
The GPU's own memory. The model's weights and the intermediate values of training must fit in it, so it limits how large a model and batch one GPU can handle.
Why do AI companies need so many chips?
Models have billions of parameters and are trained on vast datasets. One chip has neither the memory nor the speed, so the work is divided across thousands.
Conclusion
AI training is matrix arithmetic at an almost unimaginable scale, and that arithmetic happens to be perfectly parallel. GPUs, built to paint millions of pixels at once, turned out to be the right tool. The number of chips needed is a direct consequence of model size, data size and the wish to finish in weeks instead of centuries.
Related articles
- How Neural Networks Actually Learn
- What Is Gradient Descent? An Intuitive Explanation
- Transformers Explained: The Architecture Behind Modern AI
- How the CPU Cache Makes Code Fast (and How to Write Cache-Friendly Code)
