The short answer
Quick answer: Overfitting happens when a machine learning model learns its training data too closely, including the noise, quirks and coincidences, instead of the general pattern underneath. It scores brilliantly on the examples it was trained on and poorly on new ones. It is like a student who memorises last year's exam answers: perfect on those exact questions, lost when the questions change. The tell-tale sign is a growing gap between performance on training data and on held-out data. The remedies are more data, a simpler model, regularisation, and stopping training before memorisation sets in.
The whole point is generalisation
A model is useful only if it works on data it has not seen. Doing well on the training set is a means, not the goal. The ability to perform on new data is called generalisation. For how a model is fitted to its training data in the first place, see how neural networks learn.
Every dataset is a mix of:
- Signal: the real, repeatable pattern.
- Noise: measurement error, mislabelled examples and chance coincidences.
A model with enough capacity can fit both. Fitting the noise is overfitting.
Too simple, too complex, about right
Imagine fitting a curve through scattered points that roughly follow a gentle bend.
| Underfitting | Good fit | Overfitting | |
|---|---|---|---|
| The curve | A straight line that misses the bend | A smooth curve through the trend | A wild squiggle through every point |
| Training error | High | Low | Very low, near zero |
| Error on new data | High | Low | High |
| Diagnosis | Too simple | Balanced | Too flexible for the data available |
This is often described as the bias-variance trade-off:
- Bias is error from assumptions that are too simple. The model cannot capture the pattern.
- Variance is error from being too sensitive. Train on a slightly different sample and you get a very different model.
Underfitting is high bias. Overfitting is high variance.
How to detect it
You cannot see overfitting by looking at training results alone. You need data the model has never trained on.
Split the data
| Set | Used for |
|---|---|
| Training set | Fitting the model |
| Validation set | Choosing settings and deciding when to stop |
| Test set | A final, one-time estimate of real-world performance |
Watch the curves
During training, plot the loss on the training set and on the validation set.
- Early on, both fall. The model is learning real patterns.
- Then the training loss keeps falling, while the validation loss levels off.
- Eventually the validation loss starts to rise. From here, the model is memorising.
The widening gap between the two lines is overfitting. Google's Machine Learning Crash Course shows these curves with interactive examples.
For small datasets, cross-validation gives a steadier estimate: split the data into several parts, train several times with a different part held out each time, and average the results.
What causes it
- Too little data for the complexity of the model.
- Too much capacity: many parameters relative to the number of examples.
- Training for too long.
- Noisy or mislabelled data.
- Too many features, some of which correlate with the answer by pure chance.
- Unrepresentative data. The training set differs from what the model will meet in use.
How to prevent it
Get more data
The most reliable cure. With more examples, coincidences average out and real patterns stand out.
Augment the data
Create extra training examples by modifying existing ones in ways that do not change the answer: flipping, cropping or recolouring images; adding background noise to audio; rephrasing text.
Simplify the model
Fewer layers, fewer parameters, fewer input features, or a shallower decision tree.
Regularise
Regularisation discourages complexity during training.
| Technique | How it works |
|---|---|
| L2 regularisation (weight decay) | Adds a penalty for large weights, keeping the model smooth |
| L1 regularisation | Pushes many weights to exactly zero, effectively selecting features |
| Dropout | Randomly switches off a fraction of neurons at each training step, so the network cannot rely on any single path |
| Early stopping | Halt training when validation performance stops improving, and keep the best version |
Early stopping is the simplest and most widely used. It follows directly from the curves described above. The training loop itself is covered in gradient descent explained.
Combine models
Ensembles average the predictions of several models. Each overfits in a different way, and their errors partly cancel. Random forests are built on this idea.
Start from a pre-trained model
Transfer learning takes a model already trained on a huge general dataset and adapts it to your small one. It needs far less data than training from scratch.
Data leakage: the hidden version
Sometimes the validation score looks excellent because the model has, in effect, seen the answers.
- The same or near-identical examples appear in both training and test sets.
- A feature gives the answer away. A model predicting which patients have an illness uses a "treatment prescribed" column that is only filled in after diagnosis.
- Information from the future is used to predict the past in time-based data.
- Preprocessing was fitted on the whole dataset before splitting.
The model looks superb in testing and fails in production. Whenever results seem too good, suspect leakage. Split data before any preprocessing, and split time series by time.
A related trap is overfitting to the test set. If you repeatedly adjust your model after checking the test score, you gradually tune it to that particular test set. Keep the final test set untouched until the end.
Large models complicate the picture
Classical theory says that making a model bigger eventually makes overfitting worse. Modern deep learning produced a surprise: very large networks, with far more parameters than training examples, often generalise better. Test error can fall, rise, and then fall again as models grow, a pattern known as double descent.
The reasons are still being studied. The practical rules remain unchanged: hold data out, watch the gap, and regularise.
Large language models show a version of this problem too. They can reproduce passages that appeared many times in their training data, and scores on public benchmarks become unreliable if the benchmark questions leaked into the training set. How memorisation relates to made-up answers is discussed in why LLMs hallucinate.
A short checklist
- Always keep a validation set and a separate test set.
- Plot training and validation curves.
- Start with a simple model and add complexity only if it underfits.
- Use early stopping and weight decay by default.
- Get more, and more varied, data if you can.
- Be suspicious of results that look too good.
- Monitor performance after deployment, because real-world data drifts.
Frequently asked questions
What is overfitting in simple terms?
When a model memorises its training examples, including their random quirks, and so performs badly on new data.
What is the difference between overfitting and underfitting?
An overfitted model is too closely tailored to the training data. An underfitted model is too simple to capture the pattern at all. Both perform poorly on new data.
How do I know if my model is overfitting?
Compare training and validation performance. A large gap, or validation loss rising while training loss falls, indicates overfitting.
Does more data always fix overfitting?
It usually helps, provided the new data is relevant and varied. More copies of the same biased or mislabelled data will not.
Conclusion
Overfitting is the central risk in machine learning: a model that looks excellent on the data you have and fails on the data you care about. The defence is discipline more than cleverness. Keep honest held-out data, watch the gap between training and validation, prefer simpler models and more data, and stop training when the validation score stops improving.
Related articles
- How Neural Networks Actually Learn
- What Is Gradient Descent? An Intuitive Explanation
- How Recommendation Systems Know What You Want to Watch
- Why LLMs Hallucinate
