The previous chapter ended on a limit: demonstrations teach a model the shape of an answer, not how to derive one. Show a model three solved arithmetic problems and it learns the format; it does not thereby learn to calculate. For anything needing several dependent steps — multi-step arithmetic, logical deduction, questions whose answer requires combining two separate facts — few-shot prompting often leaves accuracy roughly where it found it.
Chain-of-thought (CoT) prompting, introduced in 2022, addresses that directly: rather than asking for the answer, ask the model to work through the problem and state the answer at the end.
Why a model fails at a problem it "knows"
Take a question a ten-year-old can handle:
The cafeteria had 23 apples. They used 20 for lunch, then bought 6 more. How many apples do they have now?
Asked for the answer alone, models have historically produced things like 49 — the result of adding all three numbers and ignoring that one of them is a subtraction. The arithmetic is trivial. The failure is not about arithmetic.
The explanation that comes to mind first — insufficient training on maths — is wrong, and the real one is more useful. A model produces one token at a time, and each token gets a fixed amount of computation. Answering directly means performing the whole procedure — parse the sequence of events, subtract, hold the intermediate result, add — inside the single forward pass that emits one number. There is nowhere to put the intermediate value.
Generated tokens are the missing storage. Once the model writes "23 − 20 = 3", that result is in the context, and the next step can attend to it like any other text. Reasoning becomes a sequence of easy steps, each with a full forward pass of its own, instead of one hard step.
This is why the technique helps even on problems the model has demonstrably seen. It is not supplying knowledge; it is supplying working memory. Without it the model is doing mental arithmetic; with it, it is showing its work.
Prompted to reason first, the same model produces:
The cafeteria started with 23 apples. They used 20, so 23 − 20 = 3 remain. They bought 6 more, so 3 + 6 = 9. The answer is 9.
Two ways to trigger it
Zero-shot CoT adds an instruction and nothing more. Appending "Let's think step by step" to the question is enough to shift the model into producing a reasoning trace before its answer. The effect is out of proportion to the effort: in the study that established it, accuracy on a grade-school maths benchmark rose from roughly 10% to roughly 41% for the model tested, from one added sentence.
Why that phrasing specifically is not fully understood. Many variants were tried — "let's be systematic", "let me think about this carefully", and others — and this one performed most consistently. The likeliest explanation is unglamorous: text following that phrase in the training data tends to be careful, worked-through explanation, so the phrase selects for that register.
Few-shot CoT shows the reasoning instead of requesting it. Supply two or three examples in which each answer is preceded by its full working, and the model continues the pattern on the new question. It takes more context, and it tends to work better on harder or more specialised problems, because the examples fix how to reason and not merely that reasoning should happen — which style of decomposition, what to show, how much detail.
The distinction mirrors teaching. Zero-shot CoT tells a student to show their work. Few-shot CoT works two problems on the board first.
| Zero-shot CoT | Few-shot CoT | |
|---|---|---|
| What you supply | One instruction | Worked examples |
| Context cost | Negligible | Several examples |
| Fixes | That the model reasons | How the model reasons |
| Best for | Straightforward multi-step problems | Harder or domain-specific reasoning |
Start with zero-shot. It is one sentence, and if it is enough you have spent nothing.
Making the reasoning more robust
A single chain is a single path, and a model can take a wrong turn early and continue fluently from it. Two refinements address this.
Self-consistency accepts that one chain is a sample, not a verdict. Instead of generating once, sample the model several times with enough randomness that the reasoning paths genuinely differ — five to ten is typical — extract the final answer from each, and take the majority. Different valid routes to a correct answer tend to converge on it; errors tend to scatter. Taking the mode filters out the chains that went wrong, and the accuracy gain over a single chain is substantial. The cost is proportional: ten samples is ten times the inference.
Structuring the reasoning helps where the task is deductive rather than numerical. Prompting the model to reason in explicit conditional form — stating premises, then the rules applied, then what follows — constrains it more tightly than free-form narration. Natural-language reasoning is easy to make sound valid; an explicit structure makes a missing step more visible. This matters for rule-based problems, eligibility and policy questions, and anything where a plausible-sounding paragraph can hide a logical gap.
What this means for a retrieval system
Two things follow for RAG specifically.
Reasoning tokens compete with retrieved passages. A chain of thought can run to hundreds of tokens, and self-consistency multiplies both cost and latency. In a pipeline already spending context on retrieved documents, that is a real budget decision: reason only where the question needs it, rather than on every request.
A reasoning trace is not a guarantee. A model can produce a clean, logical-looking chain and still reach a wrong answer — and worse, it can produce a chain that is not what actually drove the answer at all. Explanation and computation are different things, and the gap between them is covered in Model Interpretability. For a retrieval system the practical consequence is familiar: a reasoning trace, like a fluent answer, is only trustworthy to the extent it is grounded in retrieved sources you can check.
Chain-of-thought becomes most useful in this course once questions require combining several retrieved passages rather than reading one. Those multi-hop cases are exactly where a single compressed step fails, and they return in Agentic RAG.