The short answer
Quick answer: Gradient descent is a method for finding the settings that make an error as small as possible. Start with a guess. Measure which direction makes the error increase fastest (the gradient). Take a small step in the opposite direction. Repeat until the error stops falling. It is like walking downhill in fog: you cannot see the valley, but you can feel the slope under your feet and keep stepping down. Almost every machine learning model, from a simple straight-line fit to the largest language models, is trained this way.
The problem it solves
A model has adjustable numbers called parameters (or weights). A loss function measures how wrong the model's predictions are for a given set of parameters. Training means finding the parameters that make the loss as low as possible.
For a tiny model you might solve this with algebra. For a neural network with billions of parameters, there is no formula. You need a procedure that improves the parameters gradually. That is gradient descent.
The hill in the fog
Picture the loss as a landscape. Every possible setting of the parameters is a location, and the height at that location is the loss. You want the lowest point.
You are standing somewhere at random, in thick fog. You cannot see the whole landscape, but you can tell which way the ground slopes where you stand. So:
- Feel for the steepest downhill direction.
- Take a step that way.
- Repeat.
Eventually you reach a place where the ground is flat in every direction: a low point.
The actual rule
The gradient is a list of numbers, one per parameter, saying how much the loss would change if that parameter increased slightly. It points uphill. So we step the other way:
parameter = parameter - learning_rate × gradient
A tiny example. Suppose the loss is (w - 3)². The lowest point is obviously at w = 3. The slope at any w is 2 × (w - 3).
| Step | w | Slope | Update with learning rate 0.1 |
|---|---|---|---|
| 0 | 0.00 | -6.00 | 0.00 + 0.60 = 0.60 |
| 1 | 0.60 | -4.80 | 0.60 + 0.48 = 1.08 |
| 2 | 1.08 | -3.84 | 1.08 + 0.38 = 1.46 |
| 3 | 1.46 | -3.07 | 1.46 + 0.31 = 1.77 |
| ... | ... | ... | ... |
| 30 | 2.996 | -0.007 | About 3.00 |
Notice the steps shrink as the slope flattens. The algorithm slows down by itself as it nears the bottom.
In a neural network the same thing happens for every parameter at once. The gradients are computed by backpropagation; see how neural networks actually learn. Google's Machine Learning Crash Course has interactive exercises on this.
The learning rate
The learning rate is the step size, and it is the most important setting in training.
- Too small: progress is painfully slow, and training may stall before reaching a good solution.
- Too large: each step overshoots the valley. The loss bounces around or grows without limit.
- About right: steady decrease.
In practice the learning rate changes during training. A common schedule starts with a short warm-up from a small value, then gradually decays, so the model takes bold steps early and careful ones at the end.
Batch, stochastic and mini-batch
To compute the exact gradient you would need to run the model over the entire training set. With billions of examples, that is impractical for a single step.
| Variant | Data used per step | Character |
|---|---|---|
| Batch gradient descent | The whole dataset | Accurate, very slow, needs huge memory |
| Stochastic gradient descent (SGD) | One example | Fast, very noisy |
| Mini-batch gradient descent | A small batch, such as 32 to several thousand examples | The standard: a good estimate at a fraction of the cost |
A mini-batch gives a rough estimate of the true gradient. It is noisy, like walking downhill while being jostled. Surprisingly, the noise helps: it shakes the model out of poor spots and tends to lead to solutions that work better on new data. Mini-batches also suit GPUs, which process a whole batch in parallel.
One pass through the entire dataset is called an epoch.
Getting stuck
The landscape of a real model is not a smooth bowl.
- Local minima. A dip that is not the lowest point overall.
- Saddle points. Flat in some directions, downhill in others, like a mountain pass. In models with millions of dimensions these are far more common than true local minima.
- Plateaus. Long flat regions where the gradient is tiny and progress crawls.
- Ravines. Steep in one direction and shallow in another, causing zigzagging.
A reassuring finding from deep learning practice: in very high-dimensional models, most of the low points that training reaches are roughly equally good. The goal is a good solution, not the single best one.
Better optimisers
Plain gradient descent has been improved in several ways.
Momentum. Keep a running average of recent gradients and move along that. Like a ball rolling downhill, it builds speed in a consistent direction and smooths out zigzags.
Adaptive learning rates. Give each parameter its own step size, based on how large its gradients have been. Parameters that rarely change get larger steps.
Adam. Combines both ideas. Introduced in the paper Adam: A Method for Stochastic Optimization, it works well with little tuning and is the default choice for most deep learning. A variant called AdamW is the standard for training transformers.
| Optimiser | Idea | Notes |
|---|---|---|
| SGD | Step against the gradient | Simple; needs careful tuning |
| SGD with momentum | Add velocity | Still widely used, especially in vision |
| RMSProp | Per-parameter step sizes | A precursor to Adam |
| Adam / AdamW | Momentum plus per-parameter step sizes | The common default |
Things that go wrong
- Exploding gradients. Gradients become enormous and the weights turn into infinities or
NaNvalues. The fix is gradient clipping, which caps their size. For why numbers overflow like this, see floating point maths. - Vanishing gradients. In deep networks, gradients can shrink towards zero as they pass back through many layers, so early layers stop learning. Better activation functions, careful initialisation, normalisation layers and skip connections solved this in practice.
- Unscaled inputs. If one input ranges from 0 to 1 and another from 0 to 100,000, the landscape is a narrow ravine. Normalising inputs to a similar scale makes training far easier.
- Training too long. The loss on training data keeps falling while performance on new data gets worse. See overfitting explained.
It only works on smooth problems
Gradient descent needs a loss that changes smoothly when the parameters change, so that a slope exists. That is why machine learning uses smooth loss functions such as cross-entropy instead of optimising accuracy directly: accuracy jumps in steps and has no useful gradient.
Frequently asked questions
What is a gradient?
A list of slopes, one for each parameter, showing how the loss changes as that parameter changes. It points in the direction of steepest increase.
What is the difference between gradient descent and backpropagation?
Backpropagation calculates the gradients. Gradient descent uses them to update the parameters. They are two halves of training.
What is a good learning rate?
It depends on the model and optimiser. Values around 0.001 are a common starting point for Adam, but it should be tuned, usually with a warm-up and decay schedule.
Does gradient descent always find the best solution?
No. It finds a low point, not necessarily the lowest. For large neural networks, the solutions it finds are usually good enough.
Conclusion
Gradient descent is a simple idea applied at huge scale: measure the slope, step downhill, repeat. The learning rate decides how boldly to step, mini-batches make each step affordable, and optimisers like Adam make the walk smoother. Nearly everything described as "training" an AI model is this loop running billions of times.
Related articles
- How Neural Networks Actually Learn
- Overfitting Explained: When Your Model Memorizes Instead of Learns
- Why Training AI Needs So Many GPUs
- Why Floating Point Math Gives You 0.1 + 0.2 = 0.30000000000000004
