Part 5

Using SVM well is mostly about preprocessing, scaling, and tuning.

SVM can be powerful, but it is sensitive to data preparation. A clean pipeline matters as much as the algorithm.

When to use SVM

Model-choice question

What kind of dataset makes SVM a good candidate?

Good fitWhy
small to medium datasetskernel SVM training can be expensive on very large datasets
high-dimensional featureslinear SVM works well for text and sparse features
clear margin structureSVM is designed to find confident separation
nonlinear medium-sized dataRBF or polynomial kernels may help

When not to use SVM first

Practicality question

When might logistic regression, tree models, or boosting be easier to start with?

Feature scaling is essential

Scale question

If income ranges from 30,000 to 200,000 and age ranges from 20 to 70, which feature dominates distance?

SVM uses dot products and distances. Features with larger numeric scale can dominate the model.

\[ z=\frac{x-\mu}{\sigma} \]

Standardization puts features on comparable scale.

Before scaling large-scale featuresmall-scale feature Scaling prevents one feature from dominating.

Split before preprocessing

Leakage question

Should the scaler learn mean and standard deviation from the test data?

No. Fit preprocessing only on training data. The safest pattern is a pipeline.

\[ \text{split data} \rightarrow \text{fit scaler on train} \rightarrow \text{train SVM} \]
Pipeline([
  ("scaler", StandardScaler()),
  ("svm", SVC())
])

Hyperparameter tuning

Tuning question

How should we choose kernel, \(C\), and \(\gamma\)?

Use cross-validation on training data. Keep the test set for final evaluation.

HyperparameterWhat it controls
kernellinear or nonlinear decision shape
\(C\)strictness about margin violations
\(\gamma\)locality of RBF influence
class_weightcost adjustment for imbalanced classes

Evaluation

Metric question

If the positive class is rare or high-risk, is accuracy enough?

For applied SVM projects, report metrics that match the business cost.

MetricUse when
accuracyclasses are balanced and error costs are similar
precisionfalse positives are costly
recallfalse negatives are costly
F1need balance between precision and recall
confusion matrixwant to inspect error types directly

Probability estimates

Output question

Does SVM naturally output probabilities like logistic regression?

No. SVM naturally outputs a decision score, not a probability.

\[ f(x)=w^Tx+b \]

In sklearn, probability estimates require extra work:

SVC(probability=True)
This can make training slower. If probabilities are important, consider calibration and check probability quality.

Multiclass SVM

Multiclass question

SVM is naturally binary. How can it classify three or more classes?

Common strategies:

StrategyIdea
one-vs-resttrain one classifier per class against all others
one-vs-onetrain one classifier for every pair of classes

sklearn's `SVC` uses one-vs-one internally for multiclass classification.

SVM data engineering checklist

Before training

What should we check before fitting SVM?

  1. Confirm binary or multiclass target.
  2. Handle missing values.
  3. Encode categorical variables.
  4. Scale numeric features.
  5. Use stratified train-test split for classification.
  6. Put preprocessing inside a pipeline.
  7. Tune \(C\), kernel, \(\gamma\), and class weights with cross-validation.
  8. Evaluate with suitable metrics.

Final mental model

One-minute summary

How would you explain SVM to someone in one minute?

SVM finds a decision boundary with maximum margin. The support vectors are the critical points that define the margin. Soft margin allows controlled violations. Kernels allow nonlinear boundaries. In practice, scale features and tune carefully.

\[ \hat{y}=\operatorname{sign}(w^Tx+b) \qquad \text{margin width}=\frac{2}{\|w\|} \]
Previous: KernelsNext: SVR