The short answer
Quick answer: Attention is a mechanism that lets a neural network decide, for each item it is processing, which other items are most relevant, and draw information from them in proportion. In text, each word computes a relevance score against every other word, turns the scores into weights that add up to 1, and builds a new representation of itself as a weighted mix of the others. It was a breakthrough because it gave models a direct connection between any two positions, however far apart, and because all of those connections can be computed in parallel.
The problem it solved
Early neural translation systems had two parts. An encoder read the source sentence and squeezed it into one fixed-size vector. A decoder then wrote the translation from that vector alone.
That single vector was a bottleneck. A three-word sentence and a fifty-word sentence had to fit into the same space, and translation quality fell sharply as sentences got longer. The model was being asked to memorise the whole sentence before writing a word.
In 2014, Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio proposed a fix in Neural Machine Translation by Jointly Learning to Align and Translate. Instead of relying on one summary vector, let the decoder look back at all of the encoder's states each time it produces a word, and learn which ones to focus on. When writing the French word for "cat", the model attends mostly to "cat" in the English sentence.
A human translator works the same way: they glance back at the relevant part of the source, not recite it from memory.
Self-attention
In 2017, Attention Is All You Need took the idea further. Instead of attention between two sequences, apply it within one sequence: every token attends to every other token of the same text. This is self-attention, and the paper showed that a model built from it alone, with no recurrence, outperformed everything before it. That model was the transformer; see transformers explained.
Why does a word need to look at other words? Because meaning depends on context.
- "I sat by the river bank" and "I paid the cheque into the bank" use the same word for different things.
- In "The trophy didn't fit in the suitcase because it was too big", "it" means the trophy. Change "big" to "small" and "it" means the suitcase.
Self-attention lets each word's representation be reshaped by the words around it.
How it works: queries, keys and values
Each token starts as a vector (its embedding; see how embeddings work). From that vector, three new vectors are produced using learned weight matrices:
| Vector | Role | Analogy |
|---|---|---|
| Query | What this token is looking for | A search request |
| Key | What this token offers | A label on a folder |
| Value | The information this token passes on | The folder's contents |
Then, for each token:
- Score. Compare this token's query with every token's key, using a dot product. A higher result means a better match.
- Scale. Divide by the square root of the vector size, to keep the numbers manageable.
- Normalise. Apply softmax to turn the scores into weights between 0 and 1 that add up to 1.
- Blend. Take the weighted sum of all the value vectors.
The result is the token's new representation: its own content enriched with whatever it attended to.
In one line:
Attention(Q, K, V) = softmax(Q × Kᵀ / √d) × V
For "it" in the trophy sentence, the weights might come out something like:
| Token | Weight |
|---|---|
| trophy | 0.61 |
| suitcase | 0.22 |
| big | 0.09 |
| all others | 0.08 |
Nothing tells the model that pronouns refer to nouns. The query and key matrices are learned during training, and patterns like this emerge because they help predict text.
If this "compare a query against keys by similarity" sounds like search, that is no accident. It is a soft, differentiable lookup, and the same similarity idea drives vector databases.
Multi-head attention
One set of weights can capture only one kind of relationship at a time. So transformers run several attention calculations in parallel, each with its own query, key and value matrices. Each is a head.
Different heads learn different things. Researchers have found heads that track which noun a pronoun refers to, heads that link verbs to their subjects, and heads that attend to the previous or next word. Their outputs are joined together and mixed.
Masking
When a model is trained to predict the next token, it must not see the answer. A causal mask blocks attention from each position to all later positions by setting those scores to negative infinity before the softmax, which gives them a weight of zero.
Models that only need to understand text, such as BERT, leave the mask off and let every token see the whole input in both directions.
Why it was a breakthrough
| Recurrent networks | Attention | |
|---|---|---|
| Distance between two words | Many steps; information fades | One step, at any distance |
| Processing | One word at a time | All positions at once |
| Training speed on modern hardware | Slow | Fast |
| Long documents | Poor | Good, within the context limit |
- Direct long-range connections. A word at position 2 and a word at position 2,000 are linked directly.
- Parallelism. All the scores are one large matrix multiplication, exactly the kind of work GPUs are built for. This made it practical to train on vastly more data.
- Generality. Attention works on any set of vectors: words, image patches, audio frames, amino acids.
- A window into the model. Attention weights can be inspected, which offers hints about what the model is using, though they are not a full explanation of its reasoning.
The cost
Every token attends to every other token. For a sequence of n tokens that is n × n scores.
| Sequence length | Attention scores per head per layer |
|---|---|
| 1,000 | 1,000,000 |
| 10,000 | 100,000,000 |
| 100,000 | 10,000,000,000 |
This quadratic growth is the main limit on how much text a model can consider at once, and a major reason why long prompts are slower and more expensive. Research into more efficient attention, and engineering tricks such as caching, aim to soften it.
Where attention is used today
- Language models: every layer of every GPT-style model. See how LLMs predict the next word.
- Vision: images are cut into patches, and patches attend to each other.
- Image generation: cross-attention connects the text prompt to the image being drawn. See how AI image generators work.
- Speech, video and science: including models that predict protein structures.
Frequently asked questions
What is self-attention?
Attention applied within a single sequence, so each token computes how relevant every other token in the same text is and updates its representation accordingly.
What are queries, keys and values?
Three vectors derived from each token. A token's query is matched against all keys to get relevance weights, which are then used to blend the corresponding values.
What is multi-head attention?
Several attention calculations run in parallel with separate learned weights, so the model can track several types of relationship at once.
Is attention the same as a transformer?
No. Attention is the mechanism. A transformer is a full architecture that stacks attention layers with feed-forward layers, normalisation and residual connections.
Conclusion
Attention answers one question for every token: what else here should I be paying attention to? Turning that into a learnable, parallel calculation removed the memory bottleneck of earlier models and connected every part of the input to every other. Jay Alammar's Illustrated Transformer is the best visual companion if you want to see it step by step.
Related articles
- Transformers Explained: The Architecture Behind Modern AI
- How Large Language Models Predict the Next Word
- How Embeddings Turn Words Into Numbers
- What Are Vector Databases and Why Does AI Need Them?
