The short answer
Quick answer: A recommendation system predicts which items you are most likely to enjoy, using patterns in behaviour. The central technique is collaborative filtering: people who behaved like you in the past liked certain things, so you probably will too. Modern systems represent every user and every item as a vector of numbers (an embedding) placed so that users sit near the items they like. Recommending then happens in two stages: quickly retrieve a few hundred promising candidates from millions, then rank them with a more detailed model that predicts how likely you are to click, watch or buy each one.
The problem
A streaming service has tens of thousands of titles; a shop has millions of products; a video platform has billions of videos. Nobody can browse it all. The system's job is to choose a short list for each person, in a fraction of a second.
It learns from two kinds of feedback:
| Explicit feedback | Implicit feedback | |
|---|---|---|
| Examples | Star ratings, likes, thumbs down | Clicks, watch time, purchases, skips, searches |
| Volume | Scarce | Abundant |
| Quality | Clear but biased towards strong opinions | Noisy: a click is not always enjoyment |
Most systems rely mainly on implicit signals, because there is so much more of it.
Content-based filtering
Recommend items similar to ones you liked, based on the items' own attributes: genre, cast, description, price, visual style.
If you watched three space documentaries, here is a fourth.
- Strengths: works for brand-new items with no history; needs only your own data; easy to explain.
- Weaknesses: only as good as the item descriptions; tends to keep you in a narrow lane.
Collaborative filtering
Ignore what the items are. Look only at who interacted with what. This is collaborative filtering.
Imagine a giant table with a row per user and a column per item, mostly empty:
| Film A | Film B | Film C | Film D | |
|---|---|---|---|---|
| Asha | 5 | 4 | 1 | |
| Ben | 5 | 4 | ||
| Chen | 4 | 5 | ||
| Dana | 1 | 5 |
Asha and Ben both loved Film A. Ben also loved Film C, which Asha has not seen. So recommend Film C to Asha.
There are two classic forms:
- User-based: find people with similar taste and recommend what they liked.
- Item-based: find items that tend to be liked by the same people. "Customers who bought this also bought..." is item-based.
Its power is that it finds connections nobody would think to describe. Its weakness is that it needs interaction history.
Matrix factorisation and embeddings
The table above is enormous and almost entirely blank. Matrix factorisation compresses it: learn a short vector of numbers for every user and every item, such that multiplying a user's vector by an item's vector approximates the rating.
The numbers in these vectors are called latent factors. Nobody defines them, but they often end up corresponding to something recognisable: how serious or light a film is, how action-heavy, how mainstream.
This approach became famous through the Netflix Prize competition (2006 to 2009), in which teams competed to improve the company's rating predictions.
It is the same idea as embeddings in language models: place things in a space where closeness means compatibility.
Two-tower models
The modern neural version uses two networks:
- A user tower takes the user's history and context and outputs a user vector.
- An item tower takes an item's features and outputs an item vector.
They are trained so that a user's vector is close to the vectors of items that user engaged with. Unlike plain matrix factorisation, the towers can take in any feature: device, time of day, the text of a title, an image.
The two-stage pipeline
Scoring every item for every user with a large model is far too slow. So production systems use a funnel. Google's paper Deep Neural Networks for YouTube Recommendations described this structure, and it is now standard.
| Stage | From | To | Method |
|---|---|---|---|
| Candidate generation (retrieval) | Millions | Hundreds | Embed the user, then find the nearest item vectors with approximate nearest neighbour search |
| Ranking | Hundreds | Dozens | A larger model with many features predicts engagement for each candidate |
| Re-ranking | Dozens | The final list | Apply business rules, diversity and freshness |
Candidate generation usually blends several sources: items near your embedding, trending items, items related to what you just watched, new releases. The nearest-neighbour search is the same technology as a vector database.
Ranking predicts several things for each candidate, such as the probability of a click, the expected watch time, and the likelihood of a like or a share, and combines them into a score.
Re-ranking stops the list being ten near-identical items, removes things you have already seen, and enforces policy.
Google's free Recommendation Systems course covers each stage. The same funnel shapes social feeds; see how Instagram generates your feed.
What the model is optimised for matters
A system gets very good at whatever it is told to maximise.
- Optimise for clicks, and you get clickbait.
- Optimise for watch time, and you may promote content that holds attention regardless of whether people are glad they watched.
So platforms combine several goals, add signals of satisfaction such as surveys and "not interested" feedback, and impose constraints. Choosing the objective is the most consequential design decision in the whole system.
The hard problems
Cold start
- New users have no history. Systems fall back on popular items, ask for preferences at sign-up, and use context such as location and device.
- New items have no interactions. Content features place them roughly in the right area of the space, and they are deliberately shown to a small audience to gather data.
Feedback loops
The system recommends what is popular, which makes it more popular, which makes the system recommend it more. Items that never get shown never get the chance to prove themselves. This is popularity bias.
Explore vs exploit
Should the system show what it is confident you will like, or try something new to learn more about you? Pure exploitation narrows what you see over time. Good systems reserve some recommendations for exploration.
Filter bubbles
Repeatedly reinforcing existing tastes can narrow what people encounter. Systems counter it by measuring and promoting diversity and by mixing in unfamiliar material.
Position bias
People click the first item more simply because it is first. Training on raw clicks would teach the model that whatever it placed first was good. Models have to correct for this.
Freshness
Tastes shift and new items arrive constantly, so models are retrained frequently and react to what you did in the last few minutes.
Measuring success
- Offline: test on held-out historical data, with measures such as how often the items a user actually chose appear in the top results.
- Online: A/B tests on real users, measuring engagement, retention and satisfaction.
Offline gains do not always carry over, so live experiments are the real test.
Privacy and transparency
Recommendations are built from detailed records of behaviour, which raises obvious privacy concerns. Regulation in some regions now requires platforms to explain how their recommendation systems work and to offer alternatives that are not based on profiling. Many services provide controls to view, reset or switch off personalisation.
Frequently asked questions
What is collaborative filtering?
A method that recommends items based on the behaviour of many users, on the principle that people who agreed in the past will tend to agree in future.
What is the cold start problem?
The difficulty of making recommendations for new users or new items that have no interaction history yet.
How does Netflix or YouTube know what I will like?
They learn from what you and millions of others watched, searched for and skipped, represent users and titles as vectors, and rank the titles whose vectors best match yours.
Do recommendation systems create filter bubbles?
They can narrow what people see if designed only to reinforce past behaviour. Well-designed systems deliberately add variety and exploration.
Conclusion
A recommendation system turns behaviour into geometry: users and items become points, and recommending means finding the items nearest to you, then ranking them carefully. The mathematics is well understood. The difficult questions are what to optimise for and how to stop the system from simply amplifying what is already popular.
Related articles
- How Embeddings Turn Words Into Numbers
- What Are Vector Databases and Why Does AI Need Them?
- How Instagram Generates Your Feed
- How YouTube Streams Video Without Buffering
