The short answer
Quick answer: A neural network is a large mathematical function with millions or billions of adjustable numbers called weights. It learns by repetition: show it an example, let it make a prediction (the forward pass), measure how wrong the prediction is with a loss function, work out how each weight contributed to the error (backpropagation), and nudge every weight slightly in the direction that reduces the error (gradient descent). Repeat this millions of times over many examples, and the weights settle into values that produce good predictions, including on examples the network has never seen.
Nobody programs the rules. The network finds them by adjusting numbers.
What a network is made of
Neurons
An artificial neuron is a very small calculation:
- Take several input numbers.
- Multiply each by a weight.
- Add them up, plus a bias.
- Pass the result through an activation function.
output = activation(w1·x1 + w2·x2 + ... + wn·xn + bias)
A weight says how much that input matters. The bias shifts the threshold. The activation function adds a bend, such as ReLU, which outputs zero for negative values and the value itself otherwise.
That bend is essential. Without it, stacking any number of layers would collapse into a single straight-line formula, unable to model anything curved or complicated.
Layers
Neurons are arranged in layers:
- The input layer receives the data: pixel values, word codes, measurements.
- Hidden layers transform it step by step.
- The output layer gives the answer: a probability per category, a number, the next word.
"Deep" learning simply means many hidden layers. In an image network, early layers tend to respond to edges and colours, middle layers to textures and shapes, and later layers to whole objects. Nobody designs these features. They emerge from training.
The 3Blue1Brown neural network series shows this visually and is the best introduction available.
The training loop
Imagine training a network to recognise handwritten digits.
Step 1: Start with random weights
At first the weights are random, and the network's answers are nonsense.
Step 2: Forward pass
Feed in an image of a "7". The numbers flow through each layer, and the network outputs ten scores, one per digit. Perhaps it gives "3" the highest score.
Step 3: Measure the error
A loss function turns "how wrong was that?" into a single number. The right answer is "7" with probability 1; the network gave it 0.08. The loss is high. A perfect prediction would give a loss near zero.
Learning now has a precise goal: make the loss smaller.
Step 4: Backpropagation
To reduce the loss, the network needs to know, for every weight: if this weight were slightly larger, would the loss go up or down, and by how much? That quantity is the weight's gradient.
Calculating it separately for millions of weights would be hopeless. Backpropagation does it efficiently using the chain rule from calculus. It starts from the loss at the output and works backwards through the layers, passing along each layer's share of the blame. One backward pass yields the gradient for every weight.
Step 5: Update the weights
Move each weight a small step in the direction that lowers the loss:
new_weight = old_weight - learning_rate × gradient
This is gradient descent. The learning rate controls the step size. Too big and training becomes unstable; too small and it takes forever. See gradient descent explained.
Step 6: Repeat
Do this again with the next batch of examples, and the next. One full pass through the training data is an epoch. Networks typically train for many epochs, or, for very large models, over an enormous dataset seen roughly once.
Each update makes a tiny improvement. Millions of them add up to a network that reads handwriting.
A way to picture it
Imagine a mixing desk with a million dials. A sound comes out, and it is wrong. For each dial, you are told which way to turn it, and how far, to make the sound slightly better. You turn all of them a tiny amount. Then you play the next sound and do it again.
No single dial "knows" anything. The knowledge is in the combined setting of all of them.
What "learning" means here
- The network is not storing the training examples and looking them up.
- It is finding weight values that capture patterns: shapes that make a "7", word orders that make a sentence.
- The test of learning is generalisation: performing well on new data.
To check this, data is split into a training set the network learns from and a validation or test set it never trains on. If performance is excellent on training data and poor on test data, the network has memorised instead of learned. That failure is called overfitting.
Why it needs so much data and computing power
- Data. With millions of weights to set, the network needs many examples to pin them down. Too few, and it memorises.
- Computation. Each step involves multiplying huge grids of numbers (matrices), and there are millions of steps. Graphics processors do this kind of arithmetic in parallel, which is why they dominate the field. See why training AI needs so many GPUs.
Different shapes for different data
The learning procedure is the same for all of them. What varies is how the neurons are wired.
| Architecture | Idea | Typical use |
|---|---|---|
| Fully connected | Every neuron connects to every neuron in the next layer | Small tabular problems, final layers |
| Convolutional (CNN) | Small filters slide across the input, reusing weights | Images |
| Recurrent (RNN) | Processes a sequence one step at a time, carrying a memory | Older text and speech models |
| Transformer | Every element attends to every other | Language models, and increasingly everything |
See transformers explained for the architecture behind modern AI.
What neural networks are not
- Not brains. They were loosely inspired by neurons, but the resemblance is shallow.
- Not programmed with rules. No one writes "a 7 has a horizontal line at the top".
- Not guaranteed to be right. They produce confident outputs even for inputs unlike anything they were trained on.
- Not easy to interpret. Why a particular set of weights gives a particular answer is an active research area.
Frequently asked questions
What is backpropagation in simple terms?
A method for working out how much each weight in the network contributed to the error, by passing the error backwards from the output layer to the input layer.
What is a weight?
A number that sets how strongly one neuron's output influences another. Training is the process of finding good values for all the weights.
How long does training take?
From seconds for a tiny model on a laptop to weeks or months on thousands of specialised processors for the largest models.
Do neural networks understand what they learn?
They learn statistical patterns that are useful for their task. Whether that amounts to understanding is debated; in practice, judge them by how they perform on new, varied inputs.
Conclusion
A neural network learns by a loop that is simple to state: predict, measure the error, trace the error back to every weight, adjust, repeat. The remarkable part is not the loop but what emerges from running it at scale. Michael Nielsen's free book Neural Networks and Deep Learning is an excellent next step if you want to work through the maths.
Related articles
- What Is Gradient Descent? An Intuitive Explanation
- Overfitting Explained: When Your Model Memorizes Instead of Learns
- Why Training AI Needs So Many GPUs
- Transformers Explained: The Architecture Behind Modern AI
