The Supervised Learning Setup

Data, a model with parameters, a loss function and a learning algorithm: the four pieces behind almost every machine learning system.

Deep Learning- Fundamentals to Advanced Concepts

We now have a sigmoid neuron, but no way to choose its weights. Before we build the learning algorithm, we need the general setup it lives in. Nearly every supervised learning problem has the same shape.

Living with error

Go back to the data that a single perceptron could not separate. If you run the perceptron learning algorithm on it for, say, a thousand iterations, it cannot converge, but it does end somewhere: it draws some line through the middle of the data. On a messy dataset that line misclassifies a few points.

Often that is acceptable. If a model predicts how people will vote and gets 20% of individuals wrong, it still gives a useful picture of the election. Zero error is often impossible (you would need a bizarre, island-shaped boundary) and usually unnecessary.

So from now on the goal changes from "no errors" to "as few errors as possible."

The data

You are given NN examples, each a pair (xi,yi)(x_i, y_i):

  • xi∈Rnx_i \in \mathbb{R}^n is a vector of nn input features,
  • yiy_i is the output we want to predict. For now, a single real number.

The same layout fits many problems:

ProblemInputs xxOutput yy
Oil miningsalinity, pressure, density, temperature at a locationamount of oil extracted
Bank interest ratesalary, family size, past loans, repayment historyrate that suited that customer
Film box officebudget, number of star actors, genre, music directorbox office collection

In each case the past examples (existing drilling stations, existing customers, past films) give us many known (x,y)(x, y) pairs. Put them in a table and you have a supervised learning dataset.

The unknown function

We believe yy depends on xx. In symbols, y=f(x)y = f(x). If we knew ff we could plug in a new location and read off the oil, and there would be nothing left to do. But ff is unknown. All we have are examples of its behaviour.

The model

Since we cannot find ff, we assume a form for an approximation f^\hat f that has adjustable numbers in it, the parameters θ\theta (for us, the weights ww and bias bb). The assumed form is the model. Examples:

  • y^=wx+b\hat y = wx + b: a straight line (linear regression),
  • y^=11+e−(wx+b)\hat y = \dfrac{1}{1 + e^{-(wx + b)}}: a sigmoid (logistic regression),
  • a quadratic y^=ax2+bx+c\hat y = ax^2 + bx + c,
  • a deep neural network, which we will meet later. It is just a very complicated composite function with many parameters.

The choice of form is an assumption. Some assumptions fit the data better than others.

Why parameters matter

Why not simply say y^=11+e−x\hat y = \dfrac{1}{1 + e^{-x}} and stop? Because that function has nothing adjustable. It is like declaring "y=3x+2y = 3x + 2" without looking at any data. Who says 3 and 2 are right? With no parameters there is nothing to learn, and all our data would be useless.

Parameters let one function shape adapt to different data. For the sigmoid, ww and bb decide whether the S is steep or gentle, and where it is centred.

Learning as solving for parameters

When the model is y=mx+cy = mx + c and you are given two points, you can write two equations and solve for mm and cc. That is a learning algorithm, and a perfectly good one for this tiny case.

With many parameters and many data points, this stops working, and we need something more general. We have already seen one example of such an algorithm, the perceptron learning algorithm. The next lessons develop another, gradient descent.

The objective: a loss function

A learning algorithm needs a goal. For the perceptron the goal was no misclassified point. Now it is to make the predictions y^i\hat y_i close to the true yiy_i over all the training examples. We measure closeness with a loss function and ask the algorithm to make it as small as possible.

A common choice is the squared error:

L(w,b)=12∑i=1N(f^(xi)−yi)2L(w, b) = \frac{1}{2} \sum_{i=1}^{N} \left( \hat f(x_i) - y_i \right)^2

The factor 12\tfrac12 is only there to make later algebra tidier. Multiplying the loss by a constant does not change which parameters are best.

Why square the difference?

  • Signs must not cancel. Suppose one prediction is 0.50.5 too low and another 0.50.5 too high. Summing the raw differences gives 00, which looks perfect even though both predictions are off. Squaring makes every error count.
  • We could use the absolute value, which also avoids cancellation, but it has a sharp corner at zero, so it is not differentiable there. The square is smooth, and calculus will soon matter.

The learning algorithm's task is then: find the ww and bb that make LL smallest. The parameter vector can take any real values, so there are infinitely many candidates. That is why we need an algorithm and cannot just try them all.

Training, validation and test

A useful picture is studying a textbook chapter.

  • Training. You read the chapter and learn its formulas. Your goal is to make as few mistakes as possible on material you can see. This is the error on the training data.
  • Validation. The exercises at the end let you check whether you can apply the formulas to problems you did not study directly. If you do badly, you may go back and study again. This is the validation set, used to tune and decide.
  • Test. The exam is test data. You are not allowed to go back and study from it.

The loss we minimise during learning is the training loss. How well a model does on unseen data matters as much, and we return to it later.

The four pieces

Every supervised learning setup has:

  1. Data: pairs (xi,yi)(x_i, y_i).
  2. Model: an assumed form f^\hat f with parameters.
  3. Learning algorithm: a procedure for finding good parameters.
  4. Objective (loss function): the measure of how good a set of parameters is.

We have the first two and a choice for the fourth. The third is next.

EasyModel

Why is a model with no parameters useless for learning?

MediumLoss

Why is the squared error preferred to simply summing the differences between predictions and true values?

EasyGeneralisation

What is the difference between training loss and test error?