The short answer
Quick answer: A large language model (LLM) does one thing: given a sequence of text, it calculates how likely each possible next piece of text is. The input is split into tokens, converted to numbers, and passed through a stack of transformer layers. The output is a probability for every token in the vocabulary. The system picks one, appends it to the text, and runs the whole process again. A full answer is produced one token at a time in this loop. The ability to write essays, code and explanations emerges from doing next-token prediction extremely well, after training on a vast amount of text.
Step 1: Text becomes tokens
A model does not see letters or words. Text is first split into tokens: common words, parts of words, punctuation and spaces.
"Tokenization is unbelievable"
-> ["Token", "ization", " is", " un", "believ", "able"]
Each token has an ID number in a fixed vocabulary, typically tens of thousands to a couple of hundred thousand entries. See how tokenization works.
Step 2: Tokens become vectors
Each token ID is looked up in a table and replaced by a long list of numbers, called an embedding. Tokens with related meanings have similar embeddings. See how embeddings work.
Information about each token's position is added too, because word order matters and the next stage has no built-in sense of sequence.
Step 3: Transformer layers
The vectors pass through a stack of layers, often dozens. Each layer does two things:
- Attention. Every token looks at the tokens before it and pulls in the information that is relevant. In "The trophy did not fit in the suitcase because it was too big", attention is what links "it" to "trophy". See what attention is.
- A feed-forward network. Each token's vector is processed further on its own.
Layer by layer, the vector for each position becomes a richer summary of the text so far. The architecture comes from the 2017 paper Attention Is All You Need; see transformers explained and Jay Alammar's Illustrated Transformer.
Step 4: A probability for every token
The final vector at the last position is converted into a score for every token in the vocabulary. A function called softmax turns those scores into probabilities that add up to 1.
For the prompt "The capital of France is", the output might look like:
| Next token | Probability |
|---|---|
| " Paris" | 0.92 |
| " the" | 0.02 |
| " a" | 0.01 |
| " located" | 0.01 |
| Everything else | 0.04 |
Step 5: Choosing a token
The model does not always take the most likely token. How it chooses is controlled by sampling settings.
| Setting | What it does |
|---|---|
| Greedy | Always take the top token. Predictable; can be repetitive |
| Temperature | Reshapes the probabilities. Low values favour likely tokens; high values flatten the distribution and increase variety |
| Top-k | Consider only the k most likely tokens |
| Top-p (nucleus) | Consider the smallest set of tokens whose probabilities add up to p |
This is why the same prompt can produce different answers. The randomness is deliberate: always choosing the most likely token gives dull, repetitive text.
Step 6: Repeat
The chosen token is appended to the input, and the model runs again to predict the next one. This is called autoregressive generation.
"The capital of France is" -> " Paris"
"The capital of France is Paris" -> "."
"The capital of France is Paris." -> <end>
Generation stops when the model produces a special end token or reaches a length limit. This loop is why responses appear word by word, and why longer answers take longer.
To avoid redoing all the work at every step, systems cache the internal results for earlier tokens.
The context window
A model can only consider a limited number of tokens at once: its context window. Everything it "knows" about your conversation must fit inside: the system instructions, the chat history, any documents provided, and the answer being written.
The model has no memory between requests. A chat application creates the appearance of memory by sending the earlier conversation back to the model every time. When a conversation outgrows the window, older parts must be dropped or summarised.
How it was trained
Pretraining. The model is shown enormous amounts of text. For each position, it predicts the next token, and its weights are adjusted to make the correct token more likely. The method is gradient descent, the same as for any neural network; see how neural networks learn. The text supplies its own answers, so no manual labelling is needed.
To predict text well, the model has to pick up grammar, facts, styles, reasoning patterns and the structure of code, because all of those help predict what comes next. Nobody programs these in.
Fine-tuning and alignment. The pretrained model is then trained further to follow instructions, hold a conversation, and behave helpfully and safely, using curated examples and human feedback.
Training needs enormous computing resources; see why training AI needs so many GPUs.
What this explains about LLM behaviour
- Fluent but sometimes wrong. The model produces text that is plausible. Usually plausible and true coincide; sometimes they do not. See why LLMs hallucinate.
- A knowledge cut-off. The model only knows what was in its training data, unless you supply newer information in the prompt or give it search tools. See how RAG works.
- Sensitive to wording. A different prompt changes the probabilities at every step.
- Weak at exact arithmetic and counting letters. It works on tokens, not digits or characters.
- Able to "reason" by writing. Because each token is conditioned on everything before it, working through a problem step by step gives the model better material to predict from. Modern models are often given space to do this before answering.
Is it "just autocomplete"?
In mechanism, yes: it predicts the next token. In capability, the comparison undersells it. To predict the next line of a proof, a program or a legal argument well, a system has to model a great deal about the subject. How to describe what these models have learned, and whether "understanding" is the right word, is still debated. What is clear is that the single objective of next-token prediction, at sufficient scale, yields surprisingly general abilities.
Frequently asked questions
Does an LLM look up answers in a database?
No. Its knowledge is stored in its weights as statistical patterns. It can be connected to search or databases, but that is a separate system feeding text into the prompt.
Why does the same prompt give different answers?
Because the next token is sampled from a probability distribution. Lower the temperature for more consistent output.
What is a context window?
The maximum number of tokens a model can take into account at once, including both the input and the output being generated.
Do LLMs remember earlier conversations?
Not by themselves. Applications resend earlier messages, or store notes and feed them back in, to give the appearance of memory.
Conclusion
An LLM is a next-token predictor run in a loop: tokens in, probabilities out, pick one, repeat. The transformer layers in the middle are what make the predictions good. Keeping that loop in mind explains both what these models do remarkably well and where they stumble.
Related articles
- Transformers Explained: The Architecture Behind Modern AI
- How Tokenization Works and Why It Affects AI Behavior
- Why LLMs Hallucinate
- What Is Attention and Why Was It Such a Breakthrough?
