Search

How RAG (Retrieval-Augmented Generation) Works

The short answer

Quick answer: Retrieval-augmented generation (RAG) gives a language model relevant information at the moment it answers. Ahead of time, your documents are split into chunks, converted into embeddings and stored in a searchable index. When a question arrives, the system retrieves the chunks most related to it and places them in the prompt. The model then generates an answer based on that text. It is the difference between a closed-book exam and an open-book one. RAG lets a model use private or recent information without being retrained, and it reduces fabricated answers.

Why models need it

A language model on its own has three gaps:

  • A knowledge cut-off. It knows nothing after its training data ends.
  • No private data. It has never seen your documentation, tickets or contracts.
  • A tendency to invent. When it lacks information, it may produce something plausible. See why LLMs hallucinate.

You could paste everything into the prompt, but the context window is finite and you pay for every token. Retraining the model whenever a document changes is slow and expensive. RAG is the practical middle route: fetch only what is relevant, when it is needed.

The term comes from the 2020 paper Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Phase 1: Indexing (done in advance)

1. Load. Gather the sources: PDFs, web pages, wikis, tickets, database records. Extract clean text and keep metadata such as title, date, author and access permissions.

2. Chunk. Split each document into passages.

  • Too large, and a chunk covers many topics, so its embedding is vague and it wastes prompt space.
  • Too small, and it lacks the context to make sense.

Chunks of a few hundred tokens, split on natural boundaries such as headings and paragraphs, with some overlap, are a common starting point.

3. Embed. An embedding model turns each chunk into a vector that captures its meaning. See how embeddings work.

4. Store. Save the vectors, the text and the metadata in an index, usually a vector database.

Phase 2: Answering a question

1. Embed the query with the same embedding model.

2. Retrieve. Find the chunks whose vectors are closest to the query's. Take the top handful, say 5 to 20.

3. Rerank (optional but valuable). A slower, more accurate model scores each candidate against the question and reorders them. Retrieve broadly, then keep the best few.

4. Build the prompt. Combine instructions, the retrieved chunks and the question.

Answer the question using only the context below.
If the context does not contain the answer, say you don't know.
Cite the source of each claim.

<context>
[1] Refund policy (updated March): Customers may return items within 30 days...
[2] Shipping FAQ: Return shipping is free for faulty items...
</context>

Question: Can I return a faulty kettle after three weeks?

5. Generate. The model reads the context and writes the answer.

6. Cite. Return the sources so the user can check them.

Why keyword search still matters

Embedding search finds passages with similar meaning, even when the words differ. But it can miss exact terms: a product code, an error number, a person's name.

Keyword search, typically the BM25 algorithm, is the opposite: excellent at exact matches, blind to paraphrase.

Hybrid search runs both and merges the results. It is consistently more reliable than either alone and is the usual recommendation for production systems.

Where RAG goes wrong

Most RAG failures are retrieval failures. If the right passage never reaches the model, no prompt can save the answer.

ProblemWhat happensTypical fix
Poor chunkingAn answer is split across chunks, or a chunk loses its context ("it increased by 3%": what did?)Chunk on structure; add overlap; attach context to each chunk
Vocabulary mismatchThe query and the document use different termsHybrid search; query rewriting
Vague or multi-part questionsA single search cannot cover themBreak the question down; run several searches
Too little retrievedThe key passage is missingRetrieve more, then rerank
Too much retrievedNoise distracts the model and raises costRerank; filter by metadata
Stale indexAnswers reflect old documentsRe-index on change
Model ignores the contextIt answers from memoryClear instructions; require quotations
Questions about the whole corpus"Summarise all complaints this year" is not a lookupUse aggregation or pre-built summaries, not top-k search

One well-known chunking issue is lost context. A chunk reading "The company's revenue grew 3% over the previous quarter" is useless if you cannot tell which company or quarter. Anthropic's write-up on contextual retrieval describes adding a short, chunk-specific explanation to each chunk before indexing, and combining embeddings with keyword search and reranking, which substantially reduced failed retrievals in their tests.

Improving a RAG system

  • Query rewriting. Use a model to turn a conversational question, with its pronouns and references to earlier messages, into a clear standalone search query.
  • Metadata filters. Restrict by date, product, language or department before searching.
  • Parent-document retrieval. Search over small chunks for precision, then pass the surrounding section to the model for context.
  • Multi-step retrieval. Let the model search, read, and search again when a question needs several pieces of information. This overlaps with AI agents using tools.
  • Structured data. For questions about numbers in a database, have the model write a query instead of searching text.

Measuring quality

Evaluate the two halves separately.

Retrieval

  • Did the correct passage appear in the results (recall)?
  • How high was it ranked?

Generation

  • Faithfulness: is every claim supported by the retrieved context?
  • Relevance: does it answer the question asked?
  • Citation accuracy: do the cited sources say what is claimed?

Build a test set of real questions with known answers and rerun it whenever you change chunking, models or prompts. Without measurement, changes are guesswork.

Security and access control

  • Permissions. Retrieval must respect who is allowed to see what. Filter by the user's access rights at query time; do not rely on the model to withhold information.
  • Prompt injection. A retrieved document might contain text such as "ignore previous instructions". Treat retrieved content as untrusted data, not as instructions.
  • Sensitive data. Anything indexed can surface in an answer.

RAG and its alternatives

  • Just put it all in the prompt. With large context windows, a small knowledge base can be included whole. It is simpler, and worth doing when the material fits and cost allows.
  • Fine-tuning. Teaches style and behaviour; a poor way to add facts that change.

These are compared in fine-tuning vs prompting vs RAG.

Frequently asked questions

What does RAG stand for?

Retrieval-augmented generation: retrieving relevant information and adding it to the prompt before a language model generates its answer.

Does RAG stop hallucinations?

It reduces them considerably by grounding answers in real text. It cannot prevent errors when retrieval brings back the wrong passages or the model misreads them.

Do I need a vector database for RAG?

Not strictly. Small collections can be searched in memory, and keyword search alone works for some cases. A vector index becomes useful as the collection grows.

What is chunking?

Splitting documents into smaller passages so each can be embedded, searched and inserted into a prompt independently.

Conclusion

RAG is a search system attached to a language model. The model supplies the reading and writing ability; retrieval supplies the facts. Most of the work, and most of the improvement available, is in the retrieval half: good chunks, hybrid search, reranking and honest evaluation.

Related articles

Sources and further reading

Usama Muneer

Usama Muneer

Coder, Blogger, Tech Speaker & Web Technologies Enthusiast. Passionate about working on open-source Programming languages & Tools while utilizing my Product Development skills.

Your experience on this site will be improved by allowing cookies Cookie Policy