The short answer
Quick answer: There are three main ways to adapt a language model to your needs. Prompting changes the instructions and examples you give it. RAG (retrieval-augmented generation) changes the information it has, by fetching relevant documents and adding them to the prompt. Fine-tuning changes the model itself, by training it further on your examples. Start with prompting, because it is fastest and cheapest. Add RAG when the model lacks knowledge, such as private or recent facts. Fine-tune when you need consistent behaviour, style or format that prompting cannot achieve, or a smaller, cheaper model for a narrow task. They combine well, and most production systems use more than one.
What each one changes
| Prompting | RAG | Fine-tuning | |
|---|---|---|---|
| What changes | The instructions | The information supplied | The model's weights |
| Adds new knowledge | Only what fits in the prompt | Yes, from your documents | Poorly; not reliable for facts |
| Changes behaviour or style | To a degree | No | Yes |
| Set-up effort | Minutes to hours | Days to weeks | Weeks |
| Needs training data | A few examples | Documents | Hundreds to thousands of quality examples |
| Updating | Edit the prompt | Re-index documents | Retrain |
| Can cite sources | Only what is in the prompt | Yes | No |
A useful mental model: prompting is the briefing you give a capable new colleague, RAG is the reference library you hand them, and fine-tuning is months of on-the-job training.
Prompting
Everything you can do by changing the text sent to the model:
- Clear instructions with the context the model needs: the audience, the goal, the constraints.
- Examples of good input and output (few-shot prompting).
- A role and tone in a system prompt.
- Structure, such as asking for JSON or a specific layout.
- Space to reason, letting the model work through a problem before answering.
- Chaining, splitting a task into several smaller prompts.
Strengths: immediate, cheap, no infrastructure, easy to change and to test.
Limits: the model still only knows what is in its training data or your prompt. Very long prompts cost more on every request. And some behaviours are hard to pin down with instructions alone.
Modern models respond well to careful prompting, and a surprising number of problems that look like they need fine-tuning turn out to be prompt problems. How the model reads your prompt is explained in how LLMs predict the next word.
RAG
RAG fetches relevant passages from your own data and inserts them into the prompt, so the model answers from that text. See how RAG works and the original RAG paper.
Use it when the problem is missing knowledge:
- Company documentation, policies, product data.
- Information that changes often.
- Anything after the model's training cut-off.
- When you need citations, or per-user access control.
Strengths: current information without retraining; answers that can be checked against sources; fewer fabrications (see why LLMs hallucinate).
Limits: you are building a search system, and its quality caps the answer quality. It adds latency and prompt tokens. It does nothing for tone, format or task skill.
One alternative is worth checking first: if your whole knowledge base fits comfortably in the model's context window, simply include it. That avoids retrieval altogether. Anthropic's contextual retrieval article makes the same point before describing techniques for larger collections.
Fine-tuning
Fine-tuning continues training a pre-trained model on your own examples, typically pairs of input and ideal output. The model's weights shift so that it behaves that way by default. The mechanics are those of how neural networks learn.
Updating all of a large model's weights is expensive. Parameter-efficient methods avoid that. The best known is LoRA, introduced in LoRA: Low-Rank Adaptation of Large Language Models: freeze the original weights and train a small set of additional ones. The result is a compact "adapter" that is cheap to train and store.
Use it when the problem is behaviour:
- A consistent voice or house style.
- Strict output formats, reliably.
- A narrow, repetitive task, such as classifying support tickets, done at high volume.
- Specialised conventions the base model handles poorly.
- Distillation: training a small, fast, cheap model to imitate a large one on your specific task.
- Shorter prompts, because the instructions are "baked in".
Limits:
- You need a good dataset. Quality matters more than size, and assembling it is most of the work.
- It is a weak way to add facts. The model may absorb the style of your data and still get details wrong, and it cannot cite where an answer came from.
- Knowledge is frozen at training time. New information means another training run.
- A fine-tuned model can get worse at things outside its training examples.
- You take on evaluation, versioning and hosting, and the work may need redoing when a better base model appears.
How to decide
Ask what is actually going wrong.
| Symptom | Likely remedy |
|---|---|
| The model misunderstands the task | Better prompt: clearer instructions and examples |
| It does not know your facts | RAG, or put the documents in the prompt |
| It knows the facts but the tone or format is inconsistent | Better prompt and examples first; then fine-tuning |
| It needs to look things up or take actions | Tools; see how AI agents use tools |
| It is too slow or costly at scale for a narrow task | Fine-tune a smaller model |
| Answers are right but unverifiable | RAG with citations |
A sensible order of work:
- Build an evaluation set first. Collect real examples with known good answers. Without one, you cannot tell whether any change helped.
- Push prompting as far as it goes. It is the cheapest thing to change.
- Add retrieval if the failures are about missing or outdated knowledge.
- Consider fine-tuning if failures persist in behaviour, or if cost and latency demand a smaller model.
- Measure after every change.
They work together
These are not rivals. A typical production assistant uses all three ideas:
- A carefully written system prompt sets the role and rules.
- Retrieval supplies the relevant facts for each question.
- Tools provide live data and actions.
- Occasionally, a fine-tuned model handles a high-volume sub-task such as routing or classification.
Fine-tuning and RAG are complementary: fine-tuning shapes how the model responds, and retrieval decides what it knows at that moment.
Common mistakes
- Fine-tuning to add knowledge. Use retrieval for facts.
- Fine-tuning before trying a better prompt. Most teams reach for it too early.
- No evaluation. Judging by a few hand-picked examples.
- Blaming the model for a retrieval failure. Check what was actually retrieved.
- Stuffing the prompt with everything, burying the important instructions.
Frequently asked questions
Is RAG better than fine-tuning?
They solve different problems. RAG supplies knowledge; fine-tuning changes behaviour. For questions about your own documents, RAG is usually the right choice.
When should I fine-tune a model?
When prompting and retrieval are not enough to get consistent style, format or task performance, or when you need a smaller, cheaper model for a specific high-volume job.
Is prompt engineering enough?
Often, yes. With clear instructions, examples and the right context, many applications never need anything else.
How much data do I need to fine-tune?
It varies. A few hundred high-quality examples can be enough for a narrow task; more helps for complex behaviour. Quality and coverage matter more than raw volume.
Conclusion
Match the remedy to the problem. If the model does not understand the task, improve the prompt. If it lacks information, give it the information. If it needs to behave differently in a way instructions cannot fix, train it. Work in that order, measure each step, and you will usually find the simplest option gets you most of the way.
Related articles
- How RAG (Retrieval-Augmented Generation) Works
- Why LLMs Hallucinate
- How Large Language Models Predict the Next Word
- How AI Agents Use Tools to Complete Tasks
