Search

What Is Gradient Descent? An Intuitive Explanation

The short answer

Quick answer: Gradient descent is a method for finding the settings that make an error as small as possible. Start with a guess. Measure which direction makes the error increase fastest (the gradient). Take a small step in the opposite direction. Repeat until the error stops falling. It is like walking downhill in fog: you cannot see the valley, but you can feel the slope under your feet and keep stepping down. Almost every machine learning model, from a simple straight-line fit to the largest language models, is trained this way.

The problem it solves

A model has adjustable numbers called parameters (or weights). A loss function measures how wrong the model's predictions are for a given set of parameters. Training means finding the parameters that make the loss as low as possible.

For a tiny model you might solve this with algebra. For a neural network with billions of parameters, there is no formula. You need a procedure that improves the parameters gradually. That is gradient descent.

The hill in the fog

Picture the loss as a landscape. Every possible setting of the parameters is a location, and the height at that location is the loss. You want the lowest point.

You are standing somewhere at random, in thick fog. You cannot see the whole landscape, but you can tell which way the ground slopes where you stand. So:

  1. Feel for the steepest downhill direction.
  2. Take a step that way.
  3. Repeat.

Eventually you reach a place where the ground is flat in every direction: a low point.

The actual rule

The gradient is a list of numbers, one per parameter, saying how much the loss would change if that parameter increased slightly. It points uphill. So we step the other way:

parameter = parameter - learning_rate × gradient

A tiny example. Suppose the loss is (w - 3)². The lowest point is obviously at w = 3. The slope at any w is 2 × (w - 3).

StepwSlopeUpdate with learning rate 0.1
00.00-6.000.00 + 0.60 = 0.60
10.60-4.800.60 + 0.48 = 1.08
21.08-3.841.08 + 0.38 = 1.46
31.46-3.071.46 + 0.31 = 1.77
............
302.996-0.007About 3.00

Notice the steps shrink as the slope flattens. The algorithm slows down by itself as it nears the bottom.

In a neural network the same thing happens for every parameter at once. The gradients are computed by backpropagation; see how neural networks actually learn. Google's Machine Learning Crash Course has interactive exercises on this.

The learning rate

The learning rate is the step size, and it is the most important setting in training.

  • Too small: progress is painfully slow, and training may stall before reaching a good solution.
  • Too large: each step overshoots the valley. The loss bounces around or grows without limit.
  • About right: steady decrease.

In practice the learning rate changes during training. A common schedule starts with a short warm-up from a small value, then gradually decays, so the model takes bold steps early and careful ones at the end.

Batch, stochastic and mini-batch

To compute the exact gradient you would need to run the model over the entire training set. With billions of examples, that is impractical for a single step.

VariantData used per stepCharacter
Batch gradient descentThe whole datasetAccurate, very slow, needs huge memory
Stochastic gradient descent (SGD)One exampleFast, very noisy
Mini-batch gradient descentA small batch, such as 32 to several thousand examplesThe standard: a good estimate at a fraction of the cost

A mini-batch gives a rough estimate of the true gradient. It is noisy, like walking downhill while being jostled. Surprisingly, the noise helps: it shakes the model out of poor spots and tends to lead to solutions that work better on new data. Mini-batches also suit GPUs, which process a whole batch in parallel.

One pass through the entire dataset is called an epoch.

Getting stuck

The landscape of a real model is not a smooth bowl.

  • Local minima. A dip that is not the lowest point overall.
  • Saddle points. Flat in some directions, downhill in others, like a mountain pass. In models with millions of dimensions these are far more common than true local minima.
  • Plateaus. Long flat regions where the gradient is tiny and progress crawls.
  • Ravines. Steep in one direction and shallow in another, causing zigzagging.

A reassuring finding from deep learning practice: in very high-dimensional models, most of the low points that training reaches are roughly equally good. The goal is a good solution, not the single best one.

Better optimisers

Plain gradient descent has been improved in several ways.

Momentum. Keep a running average of recent gradients and move along that. Like a ball rolling downhill, it builds speed in a consistent direction and smooths out zigzags.

Adaptive learning rates. Give each parameter its own step size, based on how large its gradients have been. Parameters that rarely change get larger steps.

Adam. Combines both ideas. Introduced in the paper Adam: A Method for Stochastic Optimization, it works well with little tuning and is the default choice for most deep learning. A variant called AdamW is the standard for training transformers.

OptimiserIdeaNotes
SGDStep against the gradientSimple; needs careful tuning
SGD with momentumAdd velocityStill widely used, especially in vision
RMSPropPer-parameter step sizesA precursor to Adam
Adam / AdamWMomentum plus per-parameter step sizesThe common default

Things that go wrong

  • Exploding gradients. Gradients become enormous and the weights turn into infinities or NaN values. The fix is gradient clipping, which caps their size. For why numbers overflow like this, see floating point maths.
  • Vanishing gradients. In deep networks, gradients can shrink towards zero as they pass back through many layers, so early layers stop learning. Better activation functions, careful initialisation, normalisation layers and skip connections solved this in practice.
  • Unscaled inputs. If one input ranges from 0 to 1 and another from 0 to 100,000, the landscape is a narrow ravine. Normalising inputs to a similar scale makes training far easier.
  • Training too long. The loss on training data keeps falling while performance on new data gets worse. See overfitting explained.

It only works on smooth problems

Gradient descent needs a loss that changes smoothly when the parameters change, so that a slope exists. That is why machine learning uses smooth loss functions such as cross-entropy instead of optimising accuracy directly: accuracy jumps in steps and has no useful gradient.

Frequently asked questions

What is a gradient?

A list of slopes, one for each parameter, showing how the loss changes as that parameter changes. It points in the direction of steepest increase.

What is the difference between gradient descent and backpropagation?

Backpropagation calculates the gradients. Gradient descent uses them to update the parameters. They are two halves of training.

What is a good learning rate?

It depends on the model and optimiser. Values around 0.001 are a common starting point for Adam, but it should be tuned, usually with a warm-up and decay schedule.

Does gradient descent always find the best solution?

No. It finds a low point, not necessarily the lowest. For large neural networks, the solutions it finds are usually good enough.

Conclusion

Gradient descent is a simple idea applied at huge scale: measure the slope, step downhill, repeat. The learning rate decides how boldly to step, mini-batches make each step affordable, and optimisers like Adam make the walk smoother. Nearly everything described as "training" an AI model is this loop running billions of times.

Related articles

Sources and further reading

Usama Muneer

Usama Muneer

Coder, Blogger, Tech Speaker & Web Technologies Enthusiast. Passionate about working on open-source Programming languages & Tools while utilizing my Product Development skills.

Your experience on this site will be improved by allowing cookies Cookie Policy