The short answer
Quick answer: A transformer is a neural network design that processes a whole sequence at once and lets every element look directly at every other element through a mechanism called self-attention. It is built from a stack of identical blocks, each containing an attention layer and a small feed-forward network, wrapped in residual connections and normalisation. Introduced in 2017, it replaced older designs that read text one word at a time. Because it handles all positions in parallel and connects distant words directly, it trains efficiently on enormous datasets, and it is the basis of GPT, Claude, Gemini, BERT and most modern AI models.
What came before
Before 2017, language models were mostly recurrent neural networks (RNNs) and their refinement, LSTMs. They read a sentence one word at a time, updating a running "memory" as they went. Two problems held them back:
- They were sequential. Word 50 could not be processed until words 1 to 49 were done, so training could not be parallelised well.
- They forgot. Information from early in a long passage faded by the time the model reached the end.
An add-on called attention helped with the second problem. Then a team at Google asked: what if attention were the whole design? Their paper, Attention Is All You Need, introduced the transformer.
The pipeline, piece by piece
1. Tokens and embeddings
Text is split into tokens (see how tokenization works), and each token is replaced by a vector of numbers, its embedding (see how embeddings work). A sentence becomes a list of vectors.
2. Positional information
Attention on its own treats its input as an unordered set. "Dog bites man" and "man bites dog" would look identical. So the model adds information about each token's position to its vector. The original paper used fixed wave patterns of different frequencies; many later models learn positions or encode relative distances with rotations.
3. Self-attention
This is the core idea. For each token, the model asks: which other tokens matter for understanding this one?
Take "The animal didn't cross the street because it was too tired." To represent "it" well, the model should draw on "animal". Self-attention computes, for each token, a set of weights over all the others and builds a new vector as a weighted blend. After this layer, the vector for "it" carries information about the animal.
The mechanics, with queries, keys and values, are covered in what attention is and why it was a breakthrough.
Multi-head attention runs several attention calculations side by side. Each "head" can specialise: one might track grammatical subject and verb, another what a pronoun refers to, another nearby words. Their results are combined.
4. Feed-forward network
After attention has mixed information between positions, each position's vector passes through a small two-layer neural network, independently of the others. If attention is tokens communicating, the feed-forward layer is each token thinking about what it gathered. A large share of a model's parameters sit in these layers.
5. Residual connections and normalisation
Two details make it possible to stack many blocks:
- Residual connections. Each sub-layer's output is added to its input, instead of replacing it. Information, and gradients during training, have a direct path through the whole stack. This avoids the vanishing gradient problem described in gradient descent explained.
- Layer normalisation. Keeps the numbers in a stable range.
6. Stack the blocks
One block is attention followed by a feed-forward network. A transformer stacks many of them: 6 in the original paper, dozens in today's large models. Lower layers tend to capture surface patterns and grammar; higher layers capture more abstract relationships.
7. Output
For a language model, the final vector at each position is turned into a probability for each possible next token. See how LLMs predict the next word.
Jay Alammar's Illustrated Transformer walks through all of this with diagrams.
Three families
The original transformer had two halves, built for translation: an encoder to read the source sentence and a decoder to write the translation. Later models kept one half or both.
| Type | How attention works | Good at | Examples |
|---|---|---|---|
| Encoder-only | Every token sees the whole input, both directions | Understanding: classification, search, embeddings | BERT, RoBERTa |
| Decoder-only | Each token sees only earlier tokens | Generating text | GPT models, Claude, Llama, Gemini |
| Encoder-decoder | Encoder reads the input; decoder generates while attending to it | Transforming one sequence to another: translation, summarisation | The original transformer, T5, BART |
Decoder-only models apply a causal mask: when processing a token, attention to later tokens is blocked. That is what lets them be trained to predict the next token without seeing the answer.
Why transformers won
- Parallel training. All positions are processed together, which suits GPUs perfectly. RNNs could not do this.
- Long-range connections. Any two tokens are one attention step apart, no matter how far apart they are in the text.
- They scale. Add more layers, more data and more computing power, and performance keeps improving. That property, more than any single clever idea, drove the last several years of AI progress.
- They are general. The same design handles text, images (cut into patches), audio, video, protein sequences and more. Anything that can be turned into a sequence of tokens can be fed to a transformer.
The costs
- Quadratic attention. Every token attends to every other, so the work grows with the square of the sequence length. Doubling the input roughly quadruples the attention cost. This is the main reason context windows are limited and long prompts cost more.
- Data and compute hungry. Transformers have few built-in assumptions about their input, so they need a lot of data to learn what other designs get for free.
- Hard to interpret. We can inspect attention weights, but understanding what a model with billions of parameters is doing internally is an open research problem.
Much current engineering targets these costs: more efficient attention variants, caching work between generation steps, and reducing numeric precision (quantisation) to shrink models.
A sense of scale
The original transformer had tens of millions of parameters. Large models today have billions to hundreds of billions. The architecture has changed remarkably little: mostly it is the same block, repeated more times, made wider, and trained on far more data.
Frequently asked questions
What is a transformer in AI?
A neural network architecture that uses self-attention to process all parts of a sequence in parallel, letting each part draw on any other. It underlies most modern language and multimodal models.
What does GPT stand for?
Generative Pre-trained Transformer: a transformer, pre-trained on large amounts of text, that generates output.
What is the difference between BERT and GPT?
BERT is encoder-only and reads text in both directions, which suits understanding tasks. GPT is decoder-only and reads left to right, which suits generating text.
Are transformers only for text?
No. They are used for images, audio, video, code, biology and robotics. The input just needs to be represented as a sequence of tokens.
Conclusion
The transformer is a simple recipe repeated many times: let every token gather information from the others with attention, process it with a small network, keep everything stable with residual connections, and stack. Its real strength is that it runs in parallel and keeps getting better as it grows, which is why one architecture now sits behind almost every major AI system.
Related articles
- What Is Attention and Why Was It Such a Breakthrough?
- How Large Language Models Predict the Next Word
- How Embeddings Turn Words Into Numbers
- Why Training AI Needs So Many GPUs
