Part 1

SVM begins with a geometric choice: which separator is safest?

Before optimization, SVM is a simple visual idea: separate the classes while staying as far as possible from the closest points.

Many separating lines

First question

If all these lines classify the training data correctly, which one feels most reliable?

wider gapmany correct boundaries

SVM prefers the boundary that has the largest safety gap from nearby points.

Hyperplane equation

Boundary question

How do we write a separating line in a way that also works for many dimensions?

In two dimensions, the boundary can be written as:

\[ w_1x_1+w_2x_2+b=0 \]

In many dimensions:

\[ w^Tx+b=0 \]

The vector \(w\) controls orientation. The number \(b\) shifts the boundary.

w^T x + b = 0 f(x) > 0f(x) < 0

Prediction comes from which side of the hyperplane the point lies on.

Prediction by sign

Sign question

If \(w^Tx+b\) is positive for a point, which class should it belong to?

\[ f(x)=w^Tx+b \qquad \hat{y}=\operatorname{sign}(f(x)) \]
ScorePredictionMeaning
\(f(x)>0\)\(+1\)point is on the positive side
\(f(x)<0\)\(-1\)point is on the negative side
\(f(x)=0\)on boundarymodel is exactly undecided

Margin and support vectors

Safety question

Should a correctly classified point very close to the boundary be treated as fully safe?

SVM wants a confident separation. The closest points determine the safety gap. These closest points are support vectors.

\[ \text{support vectors} = \text{training points closest to the boundary} \]
Far-away points usually do not decide the SVM boundary. The closest points do.

SVM vs logistic regression

Comparison question

How is SVM different from logistic regression if both can learn a linear boundary?

ModelMain focusOutput style
Logistic Regressionlearn probabilities using log lossprobability-like output
SVMmaximize margin using hinge lossdecision score; probabilities need extra calibration
Previous: OverviewNext: Scratch SVM