Chain-of-Thought Prompting

Some problems do not get easier with more examples — they need working space. Asking a model to reason step by step before answering turns a single compressed guess into a sequence of small, tractable steps, and it costs nothing but tokens.

Advanced RAG Masterclass

The previous chapter ended on a limit: demonstrations teach a model the shape of an answer, not how to derive one. Show a model three solved arithmetic problems and it learns the format; it does not thereby learn to calculate. For anything needing several dependent steps — multi-step arithmetic, logical deduction, questions whose answer requires combining two separate facts — few-shot prompting often leaves accuracy roughly where it found it.

Chain-of-thought (CoT) prompting, introduced in 2022, addresses that directly: rather than asking for the answer, ask the model to work through the problem and state the answer at the end.

Why a model fails at a problem it "knows"

Take a question a ten-year-old can handle:

The cafeteria had 23 apples. They used 20 for lunch, then bought 6 more. How many apples do they have now?

Asked for the answer alone, models have historically produced things like 49 — the result of adding all three numbers and ignoring that one of them is a subtraction. The arithmetic is trivial. The failure is not about arithmetic.

The explanation that comes to mind first — insufficient training on maths — is wrong, and the real one is more useful. A model produces one token at a time, and each token gets a fixed amount of computation. Answering directly means performing the whole procedure — parse the sequence of events, subtract, hold the intermediate result, add — inside the single forward pass that emits one number. There is nowhere to put the intermediate value.

Generated tokens are the missing storage. Once the model writes "23 − 20 = 3", that result is in the context, and the next step can attend to it like any other text. Reasoning becomes a sequence of easy steps, each with a full forward pass of its own, instead of one hard step.

This is why the technique helps even on problems the model has demonstrably seen. It is not supplying knowledge; it is supplying working memory. Without it the model is doing mental arithmetic; with it, it is showing its work.

A direct prompt compressing a multi-step problem into one token prediction and failing, beside a chain-of-thought prompt where each intermediate result becomes context for the next step
Answering directly forces the whole procedure into one forward pass. Chain-of-thought spends tokens on the intermediate results, so each step is computed separately and remains available to the next.

Prompted to reason first, the same model produces:

The cafeteria started with 23 apples. They used 20, so 23 − 20 = 3 remain. They bought 6 more, so 3 + 6 = 9. The answer is 9.

Two ways to trigger it

Zero-shot CoT adds an instruction and nothing more. Appending "Let's think step by step" to the question is enough to shift the model into producing a reasoning trace before its answer. The effect is out of proportion to the effort: in the study that established it, accuracy on a grade-school maths benchmark rose from roughly 10% to roughly 41% for the model tested, from one added sentence.

Why that phrasing specifically is not fully understood. Many variants were tried — "let's be systematic", "let me think about this carefully", and others — and this one performed most consistently. The likeliest explanation is unglamorous: text following that phrase in the training data tends to be careful, worked-through explanation, so the phrase selects for that register.

Few-shot CoT shows the reasoning instead of requesting it. Supply two or three examples in which each answer is preceded by its full working, and the model continues the pattern on the new question. It takes more context, and it tends to work better on harder or more specialised problems, because the examples fix how to reason and not merely that reasoning should happen — which style of decomposition, what to show, how much detail.

The distinction mirrors teaching. Zero-shot CoT tells a student to show their work. Few-shot CoT works two problems on the board first.

Zero-shot CoTFew-shot CoT
What you supplyOne instructionWorked examples
Context costNegligibleSeveral examples
FixesThat the model reasonsHow the model reasons
Best forStraightforward multi-step problemsHarder or domain-specific reasoning

Start with zero-shot. It is one sentence, and if it is enough you have spent nothing.

Making the reasoning more robust

A single chain is a single path, and a model can take a wrong turn early and continue fluently from it. Two refinements address this.

Self-consistency accepts that one chain is a sample, not a verdict. Instead of generating once, sample the model several times with enough randomness that the reasoning paths genuinely differ — five to ten is typical — extract the final answer from each, and take the majority. Different valid routes to a correct answer tend to converge on it; errors tend to scatter. Taking the mode filters out the chains that went wrong, and the accuracy gain over a single chain is substantial. The cost is proportional: ten samples is ten times the inference.

Structuring the reasoning helps where the task is deductive rather than numerical. Prompting the model to reason in explicit conditional form — stating premises, then the rules applied, then what follows — constrains it more tightly than free-form narration. Natural-language reasoning is easy to make sound valid; an explicit structure makes a missing step more visible. This matters for rule-based problems, eligibility and policy questions, and anything where a plausible-sounding paragraph can hide a logical gap.

What this means for a retrieval system

Two things follow for RAG specifically.

Reasoning tokens compete with retrieved passages. A chain of thought can run to hundreds of tokens, and self-consistency multiplies both cost and latency. In a pipeline already spending context on retrieved documents, that is a real budget decision: reason only where the question needs it, rather than on every request.

A reasoning trace is not a guarantee. A model can produce a clean, logical-looking chain and still reach a wrong answer — and worse, it can produce a chain that is not what actually drove the answer at all. Explanation and computation are different things, and the gap between them is covered in Model Interpretability. For a retrieval system the practical consequence is familiar: a reasoning trace, like a fluent answer, is only trustworthy to the extent it is grounded in retrieved sources you can check.

Chain-of-thought becomes most useful in this course once questions require combining several retrieved passages rather than reading one. Those multi-hop cases are exactly where a single compressed step fails, and they return in Agentic RAG.

MediumPromptingCoT

Why does chain-of-thought help on a problem the model has clearly seen in training?

EasyPromptingCoT

What is the difference between zero-shot and few-shot chain-of-thought?

MediumPromptingCoT

How does self-consistency improve on a single chain of thought?

MediumCoTRAG

In a RAG pipeline, what is the cost of enabling chain-of-thought on every request?