Prompting and In-Context Learning

Few-shot examples, chain-of-thought and self-consistency — what each actually does to the model, and why 'in-context learning' involves no learning at all.

How to Crack the AI Engineer Interview

Prompting gets dismissed as the easy part and is then asked about carefully, because the mechanism underneath it is the same mechanism that makes retrieval work. A candidate who can explain why examples in a prompt change behaviour is in a much better position to discuss everything in the rest of this module.

What gets asked

Zero-shot, one-shot, few-shot. No examples, one, or several — typically three to five. More examples cover more of the task's range and fix the output format. The spectrum is continuous.

What in-context learning actually is. No weights change. No gradients. Everything happens in one forward pass: the examples are input, and the model attends to them while generating. It works because a prompt containing demonstrations makes performing the task the most plausible continuation. This explains both properties that matter — adaptation is instant because nothing is trained, and temporary because nothing is stored.

The distinction from few-shot learning in classical ML. Meta-learning trains a model to adapt from few examples, updating weights. Prompting updates nothing. Interviewers use this to check whether you understand the mechanism or just the vocabulary.

Chain-of-thought. Asking the model to reason before answering. The reason it works is not that the model becomes smarter: each token gets a fixed amount of computation, and a multi-step problem cannot fit in the single forward pass that emits one answer token. Writing the intermediate results out puts them in the context, where the next step can attend to them. The tokens are working memory. Zero-shot CoT triggers this with an instruction; few-shot CoT demonstrates the reasoning style.

Self-consistency. Sample several reasoning chains, take the majority answer. Correct reasoning converges by different routes; errors scatter. Costs a multiple of the inference.

Prompt construction. Representative examples including the awkward cases; rigid formatting, because the model infers a schema and inconsistency damages it; clear delimiters; demonstrating the style you want rather than describing it; watching for details leaking from examples into outputs.

The follow-ups that catch people

Why does chain-of-thought help on problems the model has obviously seen? Wanted: the limit is computation per token, not knowledge. It supplies working memory, not information.

Few-shot prompting or fine-tuning? Wanted: prompting for flexibility, speed, few examples and changing tasks; fine-tuning for stable high-volume work where the per-request token cost of carrying examples accumulates, where behaviour must be consistent, or where the examples exceed the context window.

Why does inconsistent formatting hurt? Wanted: the model infers a schema from the demonstrations and continues it. Varying the pattern — even "Answer:" versus "Ans:" — weakens the signal about what the schema is.

Is a reasoning trace an explanation of how the model got there? Wanted: no, and this is the sophisticated answer. The trace improves accuracy and is not a reliable account of the computation. Interpretability work shows models producing clean reasoning for answers they reached some other way.

You enabled chain-of-thought on every request in a RAG system. What did it cost? Wanted: output tokens, context the retrieved passages also need, and latency — multiplied if self-consistency is on. Reason selectively.

How it gets worded

  • "Compare prompting, retrieval and fine-tuning as ways of getting a model to do what you want."
  • "When would you retrieve rather than fine-tune?"
  • "How do you decide between changing the prompt and changing the weights?"
  • "Can you use retrieval and fine-tuning together? What does each contribute?"
  • "The documentation changes every week. Which approach, and why?"
  • "What does each option cost — per request and up front — and at what point is fine-tuning worth it?"
  • "The data is private and cannot leave the building. How does that change the choice?"
  • "Few-shot prompting fails where fine-tuning succeeds. What does that tell you about the task?"
  • "How would you know whether retrieval is helping at all?"
  • "Here is a use case. Which of the three would you use, and in what combination?"
  • "Why does chain-of-thought help on a problem the model has clearly seen before?"
  • "You turned reasoning on for every request in a retrieval system. What did it cost you?"

Reading path

  1. Few-Shot Prompting and In-Context Learning — the mechanism, the meta-learning distinction, prompt construction, and when it stops being the right tool.
  2. Chain-of-Thought Prompting — why it works, zero-shot versus few-shot, self-consistency, and the cost inside a retrieval pipeline.
  3. Self-Attention — the mechanism that makes conditioning on context possible, if that link is not yet solid.
  4. Model Interpretability — the faithfulness problem, for the follow-up above.
  5. RAG, Fine-Tuning or Long Context — where prompting sits among the alternatives.

The connection worth making

In-context learning is the same mechanism retrieval depends on. Few-shot examples condition the model toward a behaviour; retrieved passages condition it toward facts. Same machinery, different payload. Saying this unprompted shows you see prompting and RAG as one subject rather than two techniques, which is how a system designer thinks about them — and it leads naturally into the context-budget trade-off, where every example added is a passage not retrieved.