The previous chapter put a model's problem plainly: it cannot see your data and it guesses when unsure. Retrieval is one answer. Before reaching for it, it is worth understanding the simplest intervention of all — putting examples in the prompt — because it costs nothing to try, it solves a surprising share of problems, and the mechanism behind it is exactly the mechanism RAG relies on.
Three levels of prompting
Zero-shot gives the model an instruction and nothing else: "Translate the following sentence into French." It has to work out from the instruction alone what you want and in what form. Capable models do this well for common tasks, because something resembling the task was in their training data. Where it struggles is with anything unusual, ambiguous, or particular in its output format — the model may understand the request perfectly and still answer in the wrong shape.
One-shot supplies a single worked example before the real request. You show one English sentence and its French translation, then give a new English sentence. That one demonstration fixes the format and removes most of the ambiguity.
Few-shot supplies several — typically three to five. More examples cover more of the task's range: edge cases, the borderline judgements, the variations you care about. On tasks where the model is competent but imprecise, this is routinely the difference between unusable and reliable output.
The spectrum is continuous; one-shot is just few-shot with one example.
What is actually happening
Nothing is being trained. No weight changes, no gradients, no optimiser. The examples are simply part of the input, and everything happens in a single forward pass. This is called in-context learning, and the word "learning" does a lot of misleading work.
The mechanism follows from what a language model does: continue the text it is given, plausibly. Supply a prompt consisting of three question-and-answer pairs followed by a fourth question, and the most plausible continuation is an answer in the same form. You have not taught the model the task — you have arranged the input so that performing the task is the likeliest thing to come next.
Attention is what makes this work. Generating each output token, the model can attend directly to the examples, so their format, style and reasoning are available as context. The model has seen countless analogous patterns in pretraining; your examples select which of those latent patterns to activate. It behaves less like learning and more like retrieval from what the model already holds, steered by what you put in front of it.
The phrase collides with an older one. Few-shot learning in the meta-learning sense means training a model to adapt quickly from a handful of labelled examples — weights are updated. Few-shot prompting updates nothing. The distinction is not pedantic; it explains both properties that matter in practice. Because there is no training, adaptation is instant and free. Because there is no training, it is also temporary — the next request knows nothing about these examples unless you send them again.
This connects directly to the rest of the course. In-context learning is also what makes RAG possible: retrieved passages work because a model conditions on its context. Few-shot examples condition it toward a behaviour; retrieved documents condition it toward facts. Same mechanism, different payload.
Writing prompts that work
Models are sensitive to how examples are presented, and most few-shot failures are presentation failures.
Choose examples that span the task. Three near-identical easy cases teach the model only the easy case. Include the awkward ones — the borderline sentiment, the input with a missing field, the case where the right answer is "unknown". If you want the model to handle a situation, show it one.
Keep the format rigid. The model infers a schema from your examples, so inconsistency damages it. Writing Answer: in two examples and Ans: in a third measurably hurts results. Treat examples as rows in a table: identical structure, identical labels, every time. This matters most for structured outputs — classification, parsing, anything another program will read.
Delimit clearly. Separate examples from each other and from the live query, with blank lines, a marker line, or explicit labels. The model needs an unambiguous signal that the demonstrations have ended and the real input has begun.
Demonstrate the style you want. The model copies register, length and formatting from the examples, not from your adjectives. Asking for brevity works less well than showing three brief answers.
Watch for leakage. Specific names, numbers and domain details in examples can reappear in outputs where they do not belong. Use generic placeholders unless the specifics are the point.
Stay efficient. Three to five examples is usually the sweet spot; more is not reliably better and every token competes with the rest of the prompt — which matters a great deal once retrieved passages are also in there. Note too that models weight later examples more heavily, so put the most representative one last.
When it is the right tool
Few-shot prompting fits where flexibility matters more than consistency: you have only a handful of examples and could not fine-tune on them anyway; the task changes often enough that retraining would never keep up; you need per-user or per-session context folded in on the fly; or you are prototyping and want an answer this afternoon rather than next week.
The economics are stark. Fine-tuning a large model costs real money and hours of work, and leaves you with a model to host and version. Few-shot prompting costs the tokens in one request. For anything exploratory, that difference decides it.
It stops being the right tool in four situations.
Deep reasoning. Showing worked examples teaches the shape of an answer, not how to derive one. Arithmetic, multi-step logic and anything requiring actual working often fail to improve with demonstrations alone — that needs the technique in the next chapter.
Nothing persists. The examples apply to one request. Every subsequent call must carry them again, paying the tokens again. At high volume that recurring cost eventually exceeds the one-off cost of fine-tuning.
The context window is finite. Examples compete with the system prompt, the conversation, retrieved documents and the model's own output. In a RAG system this is a live trade-off: every example you add is a passage you cannot retrieve.
It depends on model capacity. In-context learning emerged with scale. A small model may not generalise from your examples at all, and no amount of prompt craft fixes that.
The practical rule for a retrieval system: prompt first, and only reach for fine-tuning when you can point to something prompting cannot do. RAG, Fine-Tuning or Long Context takes that decision apart properly.