Course Overview

The core algorithms of classical machine learning, built from first principles: PCA and kernel PCA, k-means, maximum likelihood, Gaussian mixtures and EM, linear, ridge and lasso regression, kNN and decision trees, naive Bayes, the perceptron, logistic regression, support vector machines, bagging, boosting and a first look at neural networks.

Machine Learning Techniques

Most of today's AI headlines are about large neural networks, but most of the machine learning that runs quietly inside products is older and simpler: a principal component analysis that compresses sensor data, a k-means clustering that groups customers, a logistic regression that scores loan applications, a gradient-boosted ensemble of trees that ranks search results. These classical techniques are fast, they work with modest amounts of data, and above all they can be understood. When one of them fails, you can usually say why.

This course builds those techniques from first principles. Each chapter starts from a concrete question ("what is the best line through this cloud of points?", "how do we group points when nobody tells us the groups?", "which boundary between two classes is the safest?"), turns it into a precise mathematical problem, and solves it. Along the way the same few ideas keep returning: projection, likelihood, convexity, regularisation, kernels and loss functions. By the end you will see the algorithms as members of one family rather than a list to memorise.

Who this course is for

The course is written for undergraduates and working professionals who want to understand machine learning beyond the level of library calls. It is concept-first: every algorithm comes with its derivation, its geometric picture, its failure modes and a worked example, but no code. If you can follow a derivation with vectors and derivatives, you can follow this course.

What you should already know

Three areas of mathematics carry the whole subject, and each matches one of the three things a learning algorithm must do.

  • Linear algebra, to describe structure: vectors, dot products, projections, matrices, eigenvalues and eigenvectors. Linear Algebra and Calculus is a refresher.
  • Probability and statistics, to describe uncertainty: random variables, the Gaussian distribution, Bayes' rule, expectation and variance. Probability and Statistics covers the basics.
  • Calculus and a little optimisation, to turn data into decisions: derivatives, gradients, and finding the minimum of a function.

Learn by experimenting

The Machine Learning Lab was built for this course. It has eleven experiments you can drive in the browser: turn a line through a cloud to find the principal component, step through k-means and EM, push a polynomial into overfitting and rescue it with ridge or lasso, watch kNN, trees, naive Bayes and SVMs draw their boundaries, grow a boosted ensemble, and compare every classification loss on one axis. The perceptron, gradient descent and neural-network chapters also use the Deep Learning Lab. Every chapter links to the matching experiment.

To practise the vocabulary, the daily word games have a machine learning topic: Guess the Term, Unscramble, Crossword and Match.

How the course is organised

  1. Foundations: what machine learning is, and its main paradigms.
  2. Representation learning and PCA: compression as understanding, projections, the covariance matrix and its eigenvectors, and what PCA can and cannot do.
  3. Kernels and kernel PCA: lifting data into a richer feature space without ever computing it.
  4. Clustering: k-means, why it converges, the geometry of its clusters, and how to start it and choose K.
  5. Probabilistic models: maximum likelihood and Bayesian estimation, Gaussian mixtures and the EM algorithm.
  6. Regression: least squares, its geometry, gradient descent, kernel regression, the probabilistic view, and ridge and lasso.
  7. Classification basics: binary classification, k-nearest neighbours and decision trees.
  8. Generative classifiers: generative versus discriminative models, naive Bayes and Gaussian naive Bayes.
  9. Linear classifiers: the perceptron and logistic regression.
  10. Support vector machines: the maximum margin, the dual problem, kernels and the soft margin.
  11. Ensembles, losses and neural networks: bagging and random forests, boosting, one view of every classification loss, and a first step into neural networks.

Every chapter ends with interview-style questions that check your understanding. Many of them are the kind asked in data-science and machine-learning interviews.

About the source

This course follows the Machine Learning Techniques lectures by Prof. Arun Rajkumar, from the IIT Madras B.S. Degree Programme. We have re-taught the material in our own words, merged related lectures into single chapters, added worked examples, diagrams and simulator experiments, and checked every derivation. The credit for the original sequence and its "comprehension is compression" thread belongs to the lecturer and the programme.