Part 2

Build the SVM idea from score, margin, hinge loss, and updates.

A small scratch SVM helps students see what sklearn later hides: raw score, margin value, hinge loss, gradients, and boundary movement.

Margin lines

Margin question

If the decision boundary is \(w^Tx+b=0\), why do the safety borders become \(w^Tx+b=1\) and \(w^Tx+b=-1\)?

SVM labels are \(y_i\in\{-1,+1\}\). It scales \(w\) and \(b\) so the closest correctly classified points satisfy:

\[ y_i(w^Tx_i+b)=1 \]

For positive support vectors, \(y_i=+1\), so:

\[ w^Tx_i+b=1 \]

For negative support vectors, \(y_i=-1\), so:

\[ w^Tx_i+b=-1 \]

Why margin width is \(2/\|w\|\)

Distance question

How far is \(w^Tx+b=1\) from \(w^Tx+b=0\)?

Distance from point \(x\) to the hyperplane \(w^Tx+b=0\) is:

\[ \frac{|w^Tx+b|}{\|w\|} \]

For the positive margin line, \(w^Tx+b=1\), so distance to the decision boundary is \(1/\|w\|\). The negative margin line is also \(1/\|w\|\) away.

\[ \text{margin width}=\frac{1}{\|w\|}+\frac{1}{\|w\|}=\frac{2}{\|w\|} \]
w^T x+b=1 w^T x+b=0 w^T x+b=-1 1/||w||1/||w||

The two safety borders are parallel to the decision boundary.

Hard-margin objective

Optimization question

If wide margin means large \(2/\|w\|\), what should we do to \(\|w\|\)?

Maximizing \(2/\|w\|\) is equivalent to minimizing \(\|w\|\). SVM usually minimizes the squared version:

\[ \min_{w,b}\frac{1}{2}\|w\|^2 \]

subject to:

\[ y_i(w^Tx_i+b)\ge 1 \]
The constraint says every point must be correctly classified and outside the margin.

Hinge loss

Loss question

Which points should get zero loss: all correct points, or only safely correct points?

Hinge loss gives zero loss only when the point is correctly classified with margin.

\[ L_i=\max(0,1-y_i(w^Tx_i+b)) \]
Margin valueMeaningLoss
\(y_if(x_i)\ge 1\)correct and safe0
\(0correct side, but inside marginpositive
\(y_if(x_i)\le 0\)wrong sidelarge positive
margin=10hinge loss = max(0, 1 - margin)

Scratch training objective

Training question

What should our from-scratch linear SVM minimize?

\[ J(w,b)=\frac{1}{2}\|w\|^2+C\cdot\frac{1}{m}\sum_{i=1}^{m}\max(0,1-y_i(w^Tx_i+b)) \]

The first term prefers a wide margin. The second term penalizes margin violations. The value \(C\) controls how expensive violations are.

Gradient descent update

Update question

How does a margin-violating point push the boundary?

If \(y_i(w^Tx_i+b)\ge 1\), the point has zero hinge loss and contributes only through regularization. If \(y_i(w^Tx_i+b)<1\), it pushes the model.

\[ w := w-\eta\frac{\partial J}{\partial w} \qquad b := b-\eta\frac{\partial J}{\partial b} \]
for each epoch:
  compute scores
  find margin violations
  compute hinge-loss gradient
  update w and b

What scratch SVM teaches

Before sklearn

After building it manually, what is no longer a black box?

ConceptWhat students see in code
Score\(f(x)=w^Tx+b\)
Margin value\(y_if(x_i)\)
Loss\(\max(0,1-y_if(x_i))\)
Support vectorspoints with margin near 1
Trainingboundary changes as \(w\) and \(b\) update
Previous: GeometryNext: Soft Margin