Part 6

Support Vector Regression fits a function with a tolerance tube.

SVR uses the same support-vector philosophy for regression: instead of separating classes with a margin, it fits a line or curve and ignores small errors inside an epsilon tube.

From classification to regression

Regression question

If SVM creates a safety margin for classification, what should the margin idea become for regression?

In classification, SVM wants a boundary with a wide margin between classes. In regression, SVR wants a function where most points fall inside a tolerance band.

This tolerance band is called the epsilon tube.

\[ f(x)=w^Tx+b \]
\[ \text{epsilon tube}: \quad f(x)-\epsilon \le y \le f(x)+\epsilon \]
f(x)+epsilonf(x)-epsilononly points outside the tube are penalized

SVR cares about errors outside the epsilon tube.

Epsilon-insensitive loss

Loss question

Should a prediction error of 0.01 and 0.50 be treated the same in regression?

SVR says small errors are acceptable. If the error is within \(\epsilon\), the loss is zero.

\[ L_{\epsilon}(y,f(x))=\max(0, |y-f(x)|-\epsilon) \]
Error sizeSVR lossMeaning
\(|y-f(x)|\le\epsilon\)0prediction is good enough
\(|y-f(x)|>\epsilon\)\(|y-f(x)|-\epsilon\)only extra error beyond tolerance is penalized

SVR objective

Optimization question

What does SVR balance: flatness of the function or prediction errors outside the tube?

For linear SVR, the objective has the same spirit as SVM classification.

\[ \min_{w,b}\frac{1}{2}\|w\|^2 + C\sum_i(\xi_i+\xi_i^*) \]

Subject to:

\[ y_i-w^Tx_i-b \le \epsilon + \xi_i \]
\[ w^Tx_i+b-y_i \le \epsilon + \xi_i^* \]
\[ \xi_i,\xi_i^*\ge 0 \]
TermMeaning
\(\frac{1}{2}\|w\|^2\)keeps the regression function flat/simple
\(\epsilon\)width of the no-penalty tube
\(\xi_i,\xi_i^*\)errors above or below the tube
\(C\)cost of errors outside the tube

Support vectors in regression

Support-vector question

In classification, support vectors touch the margin. In regression, which points become support vectors?

In SVR, support vectors are the points on or outside the epsilon tube. Points well inside the tube usually do not affect the final function much.

This is the regression version of the same idea: only critical points support the solution.

Meaning of C and epsilon

Tuning question

What happens when the tube is very wide or very narrow?

ParameterSmall valueLarge value
\(\epsilon\)narrow tube, more points penalized, more sensitivewide tube, fewer points penalized, smoother fit
\(C\)more tolerant of outside-tube errorsstricter about outside-tube errors
A very small \(\epsilon\) and very large \(C\) can overfit noise. Tune both with validation.

Kernel SVR

Nonlinear regression question

What if the target changes nonlinearly with the input?

SVR can use the same kernels as SVM classification.

KernelRegression behavior
linearfits a straight regression function
polynomialfits polynomial-like curves
RBFfits smooth nonlinear curves using local similarity
\[ K(x_i,x_j)=\exp(-\gamma\|x_i-x_j\|^2) \]

For RBF SVR, \(\gamma\) again controls locality. Large \(\gamma\) can create a wiggly regression curve.

Working example idea

Example question

How would SVR fit a noisy nonlinear curve?

For a teaching notebook or live coding, use a small synthetic regression dataset:

X = sorted random values from 0 to 6
y = sin(X) + small noise

Pipeline([
  ("scaler", StandardScaler()),
  ("svr", SVR(kernel="rbf", C=10, epsilon=0.1, gamma="scale"))
])

Then plot the fitted curve and the epsilon tube around it.

SVR data engineering workflow

Workflow question

What changes when we move from SVM classification to SVR regression?

  1. Use a regression target instead of class labels.
  2. Handle missing values and encode categorical variables.
  3. Scale features before SVR.
  4. Tune \(C\), \(\epsilon\), kernel, and \(\gamma\).
  5. Evaluate using regression metrics.
MetricMeaning
MAEaverage absolute error, easy to explain
RMSEpenalizes large errors more strongly
\(R^2\)variance explained compared with predicting the mean

SVR vs linear regression

Comparison question

When would SVR be more attractive than ordinary linear regression?

ModelLoss ideaShapeGood when
Linear Regressionsquared error for all pointslinear unless features are engineeredrelationship is simple and interpretability matters
SVRignores errors inside epsilon tubelinear or nonlinear with kernelssmall/medium data, nonlinear patterns, tolerance to small errors

When to use SVR

Model-choice question

What kind of regression problem is a good candidate for SVR?

SVR can be slow on large datasets. It is also sensitive to feature scaling and hyperparameter choices.

Final SVR mental model

One-line summary

How would you explain SVR in one sentence?

SVR fits a simple or kernel-based regression function and only worries about errors that fall outside an acceptable epsilon tube.

\[ L_{\epsilon}(y,f(x))=\max(0, |y-f(x)|-\epsilon) \]
Previous: PracticeBack to Overview