Part 4

Kernels let SVM draw nonlinear boundaries while keeping the margin idea.

When a straight line fails, SVM can compare points in a transformed feature space using a kernel function.

When linear fails

Nonlinear question

Can one straight line separate points arranged like an inner circle and outer circle?

Some datasets are not linearly separable in the original feature space. A linear SVM will underfit.

a straight line cannot isolate the inner group

Feature mapping idea

Transformation question

Could the same data become separable after adding new features?

A feature map transforms input \(x\) into a new representation \(\phi(x)\).

\[ x \rightarrow \phi(x) \]

SVM can learn a linear boundary in the transformed space:

\[ w^T\phi(x)+b=0 \]

In original space, that may look like a curve.

Kernel trick

Computation question

Do we need to explicitly create all transformed features?

No. A kernel computes the inner product in transformed space directly.

\[ K(x_i,x_j)=\phi(x_i)^T\phi(x_j) \]

This lets SVM work with rich transformations without manually constructing every transformed feature.

Polynomial kernel

Curved-boundary question

What if interactions like \(x_1^2\), \(x_2^2\), or \(x_1x_2\) help separate the classes?

\[ K(x_i,x_j)=(\gamma x_i^Tx_j+r)^d \]
ParameterMeaning
\(d\)degree of polynomial flexibility
\(\gamma\)scales the dot product
\(r\)constant term

RBF kernel

Similarity question

What if nearby points should strongly influence each other, and far points should barely influence each other?

The RBF kernel measures local similarity.

\[ K(x_i,x_j)=\exp(-\gamma\|x_i-x_j\|^2) \]
DistanceRBF similarityMeaning
smallclose to 1points strongly influence each other
largeclose to 0points barely influence each other

Role of gamma

Overfitting question

What happens if each point only influences a tiny neighborhood?

GammaBoundaryRisk
small \(\gamma\)smooth, broad influenceunderfitting
large \(\gamma\)wiggly, local influenceoverfitting
small gammamedium gammalarge gamma too smoothbalancedtoo wiggly

C and gamma together

Tuning question

Can high \(C\) and high \(\gamma\) together memorize the training data?

Yes. \(C\) controls tolerance for violations. \(\gamma\) controls locality. High values for both can produce a very flexible boundary.

Use validation or cross-validation. Do not tune SVM hyperparameters by looking at test-set performance repeatedly.
Previous: Soft MarginNext: Practice