Build the SVM idea from score, margin, hinge loss, and updates.
A small scratch SVM helps students see what sklearn later hides: raw score, margin value, hinge loss, gradients, and boundary movement.
Margin lines
Margin question
If the decision boundary is \(w^Tx+b=0\), why do the safety borders become \(w^Tx+b=1\) and \(w^Tx+b=-1\)?
SVM labels are \(y_i\in\{-1,+1\}\). It scales \(w\) and \(b\) so the closest correctly classified points satisfy:
For positive support vectors, \(y_i=+1\), so:
For negative support vectors, \(y_i=-1\), so:
Why margin width is \(2/\|w\|\)
Distance question
How far is \(w^Tx+b=1\) from \(w^Tx+b=0\)?
Distance from point \(x\) to the hyperplane \(w^Tx+b=0\) is:
For the positive margin line, \(w^Tx+b=1\), so distance to the decision boundary is \(1/\|w\|\). The negative margin line is also \(1/\|w\|\) away.
The two safety borders are parallel to the decision boundary.
Hard-margin objective
Optimization question
If wide margin means large \(2/\|w\|\), what should we do to \(\|w\|\)?
Maximizing \(2/\|w\|\) is equivalent to minimizing \(\|w\|\). SVM usually minimizes the squared version:
subject to:
Hinge loss
Loss question
Which points should get zero loss: all correct points, or only safely correct points?
Hinge loss gives zero loss only when the point is correctly classified with margin.
| Margin value | Meaning | Loss |
|---|---|---|
| \(y_if(x_i)\ge 1\) | correct and safe | 0 |
\(0| correct side, but inside margin | positive | |
| \(y_if(x_i)\le 0\) | wrong side | large positive |
Scratch training objective
Training question
What should our from-scratch linear SVM minimize?
The first term prefers a wide margin. The second term penalizes margin violations. The value \(C\) controls how expensive violations are.
Gradient descent update
Update question
How does a margin-violating point push the boundary?
If \(y_i(w^Tx_i+b)\ge 1\), the point has zero hinge loss and contributes only through regularization. If \(y_i(w^Tx_i+b)<1\), it pushes the model.
compute scores
find margin violations
compute hinge-loss gradient
update w and b
What scratch SVM teaches
Before sklearn
After building it manually, what is no longer a black box?
| Concept | What students see in code |
|---|---|
| Score | \(f(x)=w^Tx+b\) |
| Margin value | \(y_if(x_i)\) |
| Loss | \(\max(0,1-y_if(x_i))\) |
| Support vectors | points with margin near 1 |
| Training | boundary changes as \(w\) and \(b\) update |