Using SVM well is mostly about preprocessing, scaling, and tuning.
SVM can be powerful, but it is sensitive to data preparation. A clean pipeline matters as much as the algorithm.
When to use SVM
Model-choice question
What kind of dataset makes SVM a good candidate?
| Good fit | Why |
|---|---|
| small to medium datasets | kernel SVM training can be expensive on very large datasets |
| high-dimensional features | linear SVM works well for text and sparse features |
| clear margin structure | SVM is designed to find confident separation |
| nonlinear medium-sized data | RBF or polynomial kernels may help |
When not to use SVM first
Practicality question
When might logistic regression, tree models, or boosting be easier to start with?
- millions of rows and frequent retraining
- many raw categorical variables and missing values
- need for direct probability estimates
- need for very simple stakeholder explanations
- limited time for scaling and hyperparameter tuning
Feature scaling is essential
Scale question
If income ranges from 30,000 to 200,000 and age ranges from 20 to 70, which feature dominates distance?
SVM uses dot products and distances. Features with larger numeric scale can dominate the model.
Standardization puts features on comparable scale.
Split before preprocessing
Leakage question
Should the scaler learn mean and standard deviation from the test data?
No. Fit preprocessing only on training data. The safest pattern is a pipeline.
("scaler", StandardScaler()),
("svm", SVC())
])
Hyperparameter tuning
Tuning question
How should we choose kernel, \(C\), and \(\gamma\)?
Use cross-validation on training data. Keep the test set for final evaluation.
| Hyperparameter | What it controls |
|---|---|
| kernel | linear or nonlinear decision shape |
| \(C\) | strictness about margin violations |
| \(\gamma\) | locality of RBF influence |
| class_weight | cost adjustment for imbalanced classes |
Evaluation
Metric question
If the positive class is rare or high-risk, is accuracy enough?
For applied SVM projects, report metrics that match the business cost.
| Metric | Use when |
|---|---|
| accuracy | classes are balanced and error costs are similar |
| precision | false positives are costly |
| recall | false negatives are costly |
| F1 | need balance between precision and recall |
| confusion matrix | want to inspect error types directly |
Probability estimates
Output question
Does SVM naturally output probabilities like logistic regression?
No. SVM naturally outputs a decision score, not a probability.
In sklearn, probability estimates require extra work:
Multiclass SVM
Multiclass question
SVM is naturally binary. How can it classify three or more classes?
Common strategies:
| Strategy | Idea |
|---|---|
| one-vs-rest | train one classifier per class against all others |
| one-vs-one | train one classifier for every pair of classes |
sklearn's `SVC` uses one-vs-one internally for multiclass classification.
SVM data engineering checklist
Before training
What should we check before fitting SVM?
- Confirm binary or multiclass target.
- Handle missing values.
- Encode categorical variables.
- Scale numeric features.
- Use stratified train-test split for classification.
- Put preprocessing inside a pipeline.
- Tune \(C\), kernel, \(\gamma\), and class weights with cross-validation.
- Evaluate with suitable metrics.
Final mental model
One-minute summary
How would you explain SVM to someone in one minute?
SVM finds a decision boundary with maximum margin. The support vectors are the critical points that define the margin. Soft margin allows controlled violations. Kernels allow nonlinear boundaries. In practice, scale features and tune carefully.