Training Neural Networks

Backpropagation, gradient descent and the optimizers built on it, plus the practical machinery — initialisation, activations, normalisation — that makes a deep stack trainable at all.

How to Crack the AI Engineer Interview

Everything a modern model does rests on this layer, and interviews test it because it is where shallow preparation shows. You can describe a transformer without understanding gradients; you cannot explain why it trains.

What gets asked

Gradient descent. Why move against the gradient, what the learning rate controls, and the difference between batch, mini-batch and stochastic updates. Expect to discuss what goes wrong when the learning rate is too large or too small.

Backpropagation. Not the derivation, usually, but the shape of it: the chain rule applied backwards through the network, reusing each layer's gradient to compute the one below. The common follow-up is why this is efficient — because the intermediate results are shared rather than recomputed per parameter.

Optimizers. Momentum, RMSProp, Adam. What problem each solves: momentum for consistent directions and oscillation, adaptive methods for per-parameter scaling. Why Adam is the default, and what it costs — two extra state tensors per parameter, which is a real memory cost at scale and resurfaces when discussing fine-tuning.

Loss functions. Cross-entropy for classification, squared error for regression, and why the choice interacts with the output activation. Cross-entropy matters doubly here: it is the language-modelling objective, and perplexity is its exponential.

Making deep networks trainable. Vanishing and exploding gradients, what residual connections do about them, why initialisation matters, why activation functions moved from sigmoid to ReLU and its descendants, and what normalisation stabilises.

The follow-ups that catch people

Why do deep networks suffer vanishing gradients, and what fixed it? Wanted: repeated multiplication by small derivatives shrinks the signal exponentially with depth; saturating activations make it worse. Fixes: non-saturating activations, careful initialisation, normalisation, and residual connections giving gradients a direct path.

Why does Adam usually beat plain SGD, and when does it not? Wanted: per-parameter adaptive step sizes plus momentum make it robust without tuning; well-tuned SGD with momentum can generalise better on some problems, and Adam costs extra optimizer state.

What does a residual connection actually do? Wanted: both the gradient path and the representational point — each block learns an adjustment to a running representation rather than rebuilding it, which is exactly what the residual stream in a transformer is.

Your training loss falls and validation loss rises. Now what? Straight back to the regularisation material from the previous chapter.

How it gets worded

Nothing below is a new topic. They are the sentences the material above tends to arrive in.

  • "Walk me through one training step, from the forward pass to the moment a weight actually changes."
  • "What does the loss give the network that lets it improve itself?"
  • "Why should stepping against the gradient make anything better?"
  • "What is the learning rate controlling? Describe the loss curve when it is far too large, and when it is far too small."
  • "Full batch, one example at a time, or mini-batches — what separates them, and why did mini-batches become the thing everyone does?"
  • "Your training set no longer fits in memory. What changes about how an update gets computed?"
  • "Stochastic updates are noisy. When does that noise cost you, and when is it doing you a favour?"
  • "The loss has been flat for an hour. How do you work out why, and how do you get moving again?"
  • "The network is not learning at all. Talk me through how you would diagnose that."
  • "Trace the gradient of the loss back to a weight in the first layer."
  • "What goes wrong if every weight starts at the same value?"
  • "Strip out the non-linear activations. What is the deep network reduced to, and which activation would you reach for instead of a sigmoid?"
  • "Your features sit on wildly different scales and nobody normalised them. What does that do to gradient-based training?"
  • "You are scaling across many GPUs and the batch size grows with them. What do you gain, and what do you pay for it?"
  • "Batch normalisation and layer normalisation — what is each one averaging over? And what was 'internal covariate shift' supposed to mean?"

Reading path

Deep Learning covers all of this from first principles. For interview preparation, this route is enough:

  1. The Supervised Learning Setup and Output and Loss Functions — the framing and the losses.
  2. Why Move Against the Gradient — the core idea, properly motivated.
  3. Backpropagation Intuition through The Backpropagation Algorithm — the mechanism, in stages.
  4. Momentum, RMSProp and AdaDelta, Adam and AdaMax, then Choosing an Optimizer — the family and how to pick.
  5. Stochastic, Mini-batch and Batch and Scheduling the Learning Rate — the practical knobs.
  6. Initialization and Activation Functions — why a deep stack trains at all.
  7. Dropout and L2 Regularization and Weight Decay — the regularisers you will be asked to name.

The Deep Learning simulator is worth an hour here. Watching optimizers take different paths across the same loss surface makes the comparison question answerable from memory rather than recitation.

Why this chapter pays off later

Three threads run straight from here into the GenAI material.

Cross-entropy is the language-modelling objective, so everything about training loss applies directly, and perplexity is just its exponential on held-out text.

Optimizer state is a memory cost, which is the reason full fine-tuning needs roughly a dozen times the model's size in memory and why parameter-efficient methods save what they save.

Residual connections and normalisation are not transformer inventions — they are the general solution to training depth, reused. Recognising that makes the transformer block much less mysterious.