Machine Learning Foundations

The classical material still gets asked, and usually as a diagnostic: bias and variance, regularisation, generative versus discriminative, and knowing which algorithm suits which problem.

How to Crack the AI Engineer Interview

Interviews for GenAI roles spend less time here than they used to, but they have not stopped. The questions serve a purpose: someone who cannot discuss overfitting is unlikely to reason well about fine-tuning on a small dataset, and someone who has never thought about generative versus discriminative models will struggle with what a language model actually is.

What gets asked

Bias and variance. The organising idea of the whole subject. Be able to say what each is, how they trade off against model capacity, and — the follow-up that catches people — what you would actually do about a model that is clearly in one regime or the other.

Overfitting and regularisation. What it looks like on training and validation curves, and the remedies: weight penalties, early stopping, dropout, more data, augmentation, ensembling. Know how each one restricts the model, not just that it helps.

The standard algorithms. Linear and logistic regression, k-nearest neighbours, decision trees, random forests, boosting, SVMs, k-means, PCA. For each: what it assumes, when it fails, what it costs. You are rarely asked to derive them; you are frequently asked to choose between them.

Generative versus discriminative. Does the model learn the joint distribution or the decision boundary? This matters more than it looks, because it is the frame for understanding what a language model is doing when it predicts a next token.

Evaluation for classifiers. Precision, recall, F1, and when accuracy misleads. Expect a class-imbalance scenario — fraud, disease, defects — where accuracy is useless and you have to say why.

The follow-ups that catch people

You say the model is overfitting. How do you know, and what do you try first? Wanted: validation loss rising while training loss falls, then a reasoned order of remedies — more data if obtainable, then regularisation, then reduced capacity — rather than a list.

Why does L2 regularisation reduce overfitting? Wanted: it penalises large weights, keeping the fitted function smoother and less able to contort to individual points.

Random forest or gradient boosting? Wanted: bagging reduces variance by averaging independent trees; boosting reduces bias by fitting sequentially to residuals, which makes it stronger and more prone to overfitting noise. Then a recommendation.

When would you use k-means rather than a Gaussian mixture? Wanted: hard assignment and spherical clusters versus soft assignment and covariance structure, with the cost difference acknowledged.

How it gets worded

The same handful of ideas arrives under many different sentences. These are the shapes to recognise.

  • "Explain the bias-variance trade-off using a model you have actually trained."
  • "Training error is near zero, validation error is climbing. What is happening, and what do you try first?"
  • "How does L1 differ from L2 in what it does to the weights, and when would you want one over the other?"
  • "Tabular data, a hundred thousand rows, forty columns. Logistic regression, gradient boosting or a neural network — argue for one."
  • "What does a model that learns the joint distribution give you that a model learning the boundary does not?"
  • "One per cent of your labels are positive. Why is accuracy useless here, and what do you report instead?"
  • "Talk me through cross-validation. When does it give you a number you should not trust?"
  • "What does the curse of dimensionality look like in something you would actually notice?"
  • "Why does averaging several weak models beat tuning one hard?"
  • "Where does language model pretraining sit among supervised, unsupervised and self-supervised learning?"

Reading path

Everything here is covered in Machine Learning Techniques. A focused route:

  1. What Machine Learning Is and Paradigms of Machine Learning — the framing.
  2. Linear Regression and Ridge and Lasso — fitting, and the first regularisation.
  3. Binary Classification and kNN, Decision Trees, Logistic Regression.
  4. Generative and Discriminative Models and Naive Bayes — the distinction above.
  5. Maximum Margin and Soft Margin SVM — SVMs, and what the margin buys.
  6. Bagging and Random Forests and Boosting — the ensemble pair, and the question above.
  7. Loss Functions Compared — ties the methods together by what they optimise.

For bias and variance and the regularisation family specifically, the Deep Learning course treats them directly in Underfitting, Overfitting and Test Error and Bias and Variance.

If clustering and dimensionality reduction come up, Clustering and k-means, Gaussian Mixture Models and Principal Component Analysis cover them.

How much is enough

For a GenAI-focused role: be fluent on bias and variance, regularisation and evaluation metrics, and able to discuss the common algorithms sensibly without deriving them. For a general ML engineer role, go deeper — the derivations do get asked.

One connection worth carrying forward: regularisation reappears everywhere in this course. LoRA works partly because constraining updates to a low-rank subspace regularises them. Distillation's soft targets regularise the student. Dropout, weight decay and early stopping are the same instinct applied in different places. Interviewers like candidates who notice that.