Soft margin SVM accepts that real data is messy.
Hard margin is elegant, but real datasets overlap. Soft margin allows controlled violations so the model can generalize better.
Why hard margin fails
Reality question
If one noisy point appears on the wrong side, should the whole boundary twist around it?
In real data, classes can overlap. A perfectly separating boundary may not exist, or may overfit badly.
Soft margin says: allow some points to violate the margin, but charge a penalty.
Slack variables
Violation question
How can we measure how much a point violates the margin?
Soft margin introduces slack variables \(\xi_i\). A slack value measures how much point \(i\) violates the margin condition.
| Slack value | Meaning |
|---|---|
| \(\xi_i=0\) | point is correctly classified and outside margin |
| \(0<\xi_i<1\) | point is inside margin but still on correct side |
| \(\xi_i\ge 1\) | point is misclassified or exactly beyond the boundary |
Soft-margin objective
Tradeoff question
Should the model prefer a wide margin or fewer violations?
The first term wants a wide margin. The second term penalizes violations. \(C\) controls how much violations hurt.
The role of C
Strictness question
What happens if mistakes become very expensive?
| Value of \(C\) | Behavior | Risk |
|---|---|---|
| Small \(C\) | more tolerant, wider margin, allows more violations | underfitting |
| Large \(C\) | stricter, tries to reduce training violations | overfitting |
Class imbalance and class weights
Imbalance question
If one class is rarer but important, should mistakes on both classes cost the same?
SVM can use class weights so minority-class mistakes become more expensive.
In sklearn, balanced class weights are roughly inversely proportional to class frequency.