The short answer
Quick answer: Retrieval-augmented generation (RAG) gives a language model relevant information at the moment it answers. Ahead of time, your documents are split into chunks, converted into embeddings and stored in a searchable index. When a question arrives, the system retrieves the chunks most related to it and places them in the prompt. The model then generates an answer based on that text. It is the difference between a closed-book exam and an open-book one. RAG lets a model use private or recent information without being retrained, and it reduces fabricated answers.
Why models need it
A language model on its own has three gaps:
- A knowledge cut-off. It knows nothing after its training data ends.
- No private data. It has never seen your documentation, tickets or contracts.
- A tendency to invent. When it lacks information, it may produce something plausible. See why LLMs hallucinate.
You could paste everything into the prompt, but the context window is finite and you pay for every token. Retraining the model whenever a document changes is slow and expensive. RAG is the practical middle route: fetch only what is relevant, when it is needed.
The term comes from the 2020 paper Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.
Phase 1: Indexing (done in advance)
1. Load. Gather the sources: PDFs, web pages, wikis, tickets, database records. Extract clean text and keep metadata such as title, date, author and access permissions.
2. Chunk. Split each document into passages.
- Too large, and a chunk covers many topics, so its embedding is vague and it wastes prompt space.
- Too small, and it lacks the context to make sense.
Chunks of a few hundred tokens, split on natural boundaries such as headings and paragraphs, with some overlap, are a common starting point.
3. Embed. An embedding model turns each chunk into a vector that captures its meaning. See how embeddings work.
4. Store. Save the vectors, the text and the metadata in an index, usually a vector database.
Phase 2: Answering a question
1. Embed the query with the same embedding model.
2. Retrieve. Find the chunks whose vectors are closest to the query's. Take the top handful, say 5 to 20.
3. Rerank (optional but valuable). A slower, more accurate model scores each candidate against the question and reorders them. Retrieve broadly, then keep the best few.
4. Build the prompt. Combine instructions, the retrieved chunks and the question.
Answer the question using only the context below.
If the context does not contain the answer, say you don't know.
Cite the source of each claim.
<context>
[1] Refund policy (updated March): Customers may return items within 30 days...
[2] Shipping FAQ: Return shipping is free for faulty items...
</context>
Question: Can I return a faulty kettle after three weeks?
5. Generate. The model reads the context and writes the answer.
6. Cite. Return the sources so the user can check them.
Why keyword search still matters
Embedding search finds passages with similar meaning, even when the words differ. But it can miss exact terms: a product code, an error number, a person's name.
Keyword search, typically the BM25 algorithm, is the opposite: excellent at exact matches, blind to paraphrase.
Hybrid search runs both and merges the results. It is consistently more reliable than either alone and is the usual recommendation for production systems.
Where RAG goes wrong
Most RAG failures are retrieval failures. If the right passage never reaches the model, no prompt can save the answer.
| Problem | What happens | Typical fix |
|---|---|---|
| Poor chunking | An answer is split across chunks, or a chunk loses its context ("it increased by 3%": what did?) | Chunk on structure; add overlap; attach context to each chunk |
| Vocabulary mismatch | The query and the document use different terms | Hybrid search; query rewriting |
| Vague or multi-part questions | A single search cannot cover them | Break the question down; run several searches |
| Too little retrieved | The key passage is missing | Retrieve more, then rerank |
| Too much retrieved | Noise distracts the model and raises cost | Rerank; filter by metadata |
| Stale index | Answers reflect old documents | Re-index on change |
| Model ignores the context | It answers from memory | Clear instructions; require quotations |
| Questions about the whole corpus | "Summarise all complaints this year" is not a lookup | Use aggregation or pre-built summaries, not top-k search |
One well-known chunking issue is lost context. A chunk reading "The company's revenue grew 3% over the previous quarter" is useless if you cannot tell which company or quarter. Anthropic's write-up on contextual retrieval describes adding a short, chunk-specific explanation to each chunk before indexing, and combining embeddings with keyword search and reranking, which substantially reduced failed retrievals in their tests.
Improving a RAG system
- Query rewriting. Use a model to turn a conversational question, with its pronouns and references to earlier messages, into a clear standalone search query.
- Metadata filters. Restrict by date, product, language or department before searching.
- Parent-document retrieval. Search over small chunks for precision, then pass the surrounding section to the model for context.
- Multi-step retrieval. Let the model search, read, and search again when a question needs several pieces of information. This overlaps with AI agents using tools.
- Structured data. For questions about numbers in a database, have the model write a query instead of searching text.
Measuring quality
Evaluate the two halves separately.
Retrieval
- Did the correct passage appear in the results (recall)?
- How high was it ranked?
Generation
- Faithfulness: is every claim supported by the retrieved context?
- Relevance: does it answer the question asked?
- Citation accuracy: do the cited sources say what is claimed?
Build a test set of real questions with known answers and rerun it whenever you change chunking, models or prompts. Without measurement, changes are guesswork.
Security and access control
- Permissions. Retrieval must respect who is allowed to see what. Filter by the user's access rights at query time; do not rely on the model to withhold information.
- Prompt injection. A retrieved document might contain text such as "ignore previous instructions". Treat retrieved content as untrusted data, not as instructions.
- Sensitive data. Anything indexed can surface in an answer.
RAG and its alternatives
- Just put it all in the prompt. With large context windows, a small knowledge base can be included whole. It is simpler, and worth doing when the material fits and cost allows.
- Fine-tuning. Teaches style and behaviour; a poor way to add facts that change.
These are compared in fine-tuning vs prompting vs RAG.
Frequently asked questions
What does RAG stand for?
Retrieval-augmented generation: retrieving relevant information and adding it to the prompt before a language model generates its answer.
Does RAG stop hallucinations?
It reduces them considerably by grounding answers in real text. It cannot prevent errors when retrieval brings back the wrong passages or the model misreads them.
Do I need a vector database for RAG?
Not strictly. Small collections can be searched in memory, and keyword search alone works for some cases. A vector index becomes useful as the collection grows.
What is chunking?
Splitting documents into smaller passages so each can be embedded, searched and inserted into a prompt independently.
Conclusion
RAG is a search system attached to a language model. The model supplies the reading and writing ability; retrieval supplies the facts. Most of the work, and most of the improvement available, is in the retrieval half: good chunks, hybrid search, reranking and honest evaluation.
Related articles
- How Embeddings Turn Words Into Numbers
- What Are Vector Databases and Why Does AI Need Them?
- Why LLMs Hallucinate
- Fine-Tuning vs Prompting vs RAG: Which Should You Use?
