RAG, Fine-Tuning or Long Context?

There are three ways to give a model knowledge it lacks: retrieve it, train it in, or paste it all into a huge context window. They solve different problems. Choosing well depends on whether you need knowledge or behaviour, how fast the data changes, and how much of it there is.

Advanced RAG Masterclass

Before building any RAG system, ask whether you need one at all. "The model doesn't know our stuff" has three possible remedies, and teams regularly choose the wrong one. RAG is not always the answer: sometimes the right move is to change the model itself, and sometimes it is simply to paste the documents into the prompt.

The three options

RAG keeps the model unchanged and supplies knowledge at question time by retrieving it from an index.

Fine-tuning keeps the knowledge in the weights. You continue training the model on your own examples so that its parameters change. Parameter-efficient methods such as LoRA and QLoRA make this cheaper by training a small set of extra weights rather than the whole network, but it is still training.

Long context skips retrieval entirely. Modern models accept hundreds of thousands of tokens, and some accept millions. If the material fits, you can include all of it in the prompt and let the model find what it needs. Uploading a handful of files to a custom assistant works this way.

Knowledge or behaviour?

The most useful question is whether the gap is about what the model knows or how the model behaves.

Fine-tuning is excellent at behaviour: tone, format, style and domain habits. A model fine-tuned on a clinician's consultations learns to respond like a careful doctor, asking the right follow-up questions and hedging appropriately. A model fine-tuned to be encouraging stays encouraging. These are patterns repeated over thousands of examples, and that is exactly what gradient descent absorbs well.

Fine-tuning is a poor way to store specific, checkable facts. Weights blend what they learn. They do not file it away with a citation, so a fine-tuned model can still mix up "15 days" and "20 days", and it cannot tell you which document the number came from. RAG is the opposite: excellent for facts, which arrive verbatim with their source, and neutral about behaviour.

How fast does the knowledge change?

Freshness is the second axis.

  • With RAG, updating knowledge means updating the index: re-embed the changed document and upsert it (Module 6). The next question sees the new policy within minutes, and no GPU is required.
  • With fine-tuning, new knowledge means another training run, with data preparation, GPU time and evaluation. That is acceptable if the domain changes every six months. If policies change weekly, you will always be behind.
  • With long context, freshness is free, because you simply send the latest files. The cost is paid elsewhere.

How much knowledge is there?

Long context is attractive because it removes a whole system: no chunking, no embeddings, no vector database. For a few documents it is often the right call. Three HR PDFs do not need a RAG pipeline. But it has limits:

  • Cost and latency grow with every token, on every question. Re-reading 500,000 tokens to answer "what is the notice period?" is wasteful, and the attention cost and key–value cache of a long prompt are not free (see The KV Cache for why long context is expensive).
  • Attention is not uniform. Models tend to use information near the start and end of a very long prompt more reliably than information buried in the middle.
  • The corpus may simply not fit. Ten thousand documents will not fit in any context window.

The weaknesses of RAG

RAG has its own costs, and it is worth stating them plainly:

  • It is only as good as its retrieval. If the question uses no words or concepts that match the right chunk, because it is vague or uses different vocabulary, the model never sees the answer. Much of this course (hybrid search, query rewriting, agentic loops) exists to reduce this weakness.
  • It is infrastructure. An ingestion pipeline, an embedding model, a vector database, a serving backend and evaluation all need building and running.

Choosing

Decision diagram: if the need is behaviour or style, fine-tune; if it is knowledge, ask whether it fits comfortably in context — if yes use long context, if no or if it changes often or needs citations, use RAG; combinations are common
A rough decision guide. Behaviour points to fine-tuning; knowledge points to RAG, unless the corpus is small and stable enough to place in the prompt. Mature systems often combine them.
SituationBest fit
Private or confidential facts, changing often, need citationsRAG
Specific tone, format or domain behaviour (medical, legal, brand voice)Fine-tuning
A few documents, occasional use, no infrastructure wantedLong context
Large corpus and specialised behaviourRAG + fine-tuned model

The options are not exclusive. A legal assistant might use a model fine-tuned to write like a careful lawyer, RAG to pull the exact clauses, and a long context window to hold the full text of the retrieved contract sections. You will see the same layering throughout this course.

MediumArchitectureFine-tuning

A hospital wants an assistant that answers in a cautious clinical style and cites the hospital's current treatment protocols, which are revised monthly. What would you build?

MediumFine-tuning

Why is fine-tuning a poor way to teach a model your company's exact leave entitlements?

EasyLong context

Your team has five PDFs totalling 60 pages, and the assistant will be used a few times a week. Do you need RAG?