Concept 5

Multiple linear regression

Real problems rarely depend on one input. Multiple regression lets many features contribute to one prediction.

Equation

\[\hat{y} = \beta_0 + \beta_1x_1 + \beta_2x_2 + \beta_3x_3 + \cdots + \beta_px_p\]

Every feature gets its own coefficient. The interpretation is:

If x1 increases by one unit while all other features stay constant, the prediction changes by b1 units.

Feature brainstorm

If we are predicting house price, what features should we use?

Think about numerical features, categorical features, useful features, suspicious features, and features that may not be available at prediction time.

How does the multivariate model fit?

The model starts with many possible planes or hyperplanes. Each choice of coefficients creates a different prediction surface and different residuals.

\[\text{residual}_i = y_i - \hat{y}_i\]

The best fit is the surface that minimizes the total squared residuals across all rows:

\[\min_{\beta_0,\beta_1,\ldots,\beta_p}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\]

Intercept moves the surface

\(\beta_0\) shifts the whole line, plane, or hyperplane up and down without changing its tilt.

Coefficients tilt the surface

Each \(\beta_j\) controls tilt in one feature direction. Larger magnitude means the prediction changes more along that feature.

All features compete together

The model chooses coefficients jointly, not one by one. A feature's coefficient depends on the other features present.

Underfit plane Better fit Too steep

Fitting means adjusting the surface until the residual distances are as small as possible overall.

Coefficient meaning: holding other features constant

In multiple regression, each coefficient is a partial effect. It answers:

\[\text{What happens to }\hat{y}\text{ if }x_j\text{ increases by 1, while other features stay fixed?}\]

This is why coefficient interpretation is more subtle than simple regression.

QuestionSimple regressionMultiple regression
What does slope mean?How \(y\) changes as one \(x\) changes.How \(y\) changes as one \(x_j\) changes while other features stay constant.
Can coefficients change when we add features?No other features exist.Yes. Adding/removing features can change coefficients.
Why?One feature explains everything it can.Features share explanatory responsibility.
If two features carry similar information, the model may split credit between them in unstable ways. This is why multicollinearity matters.

Projection intuition

Another way to think about fitting is projection. The target vector \(y\) may not be perfectly represented by the feature columns in \(X\). Linear regression finds the closest possible prediction vector \(\hat{y}\) inside the space created by those features.

\[\hat{y}=X\beta\]
\[e = y - \hat{y}\]

At the OLS solution, the residual vector is orthogonal to the feature space:

\[X^T(y - X\beta)=0\]
actual y closest prediction y_hat feature space from X residual

The prediction is the closest point the linear model can reach using the available features.

Why this helps

If important features are missing, the feature space cannot reach the true target well. The residual remains large.

If we add useful features, the feature space expands, and the model may get closer to \(y\).

If we add noisy or redundant features, the model may become harder to interpret or overfit.

From line fitting to plane fitting

With one input feature, linear regression fits a line:

\[\hat{y} = \beta_0 + \beta_1x_1\]

With two input features, it fits a plane:

\[\hat{y} = \beta_0 + \beta_1x_1 + \beta_2x_2\]

Example: house price may depend on both area and number of bedrooms.

\[\widehat{price} = \beta_0 + \beta_1(area) + \beta_2(bedrooms)\]

The model is still linear because the coefficients are added together in a straight weighted sum.

x1: area x2: bedrooms y: price fitted plane

For two features, predictions lie on a plane. Residuals are vertical distances from points to the plane.

Advertising example

\[\widehat{sales} = \beta_0 + \beta_1(TV) + \beta_2(Radio) + \beta_3(Newspaper)\]
CoefficientInterpretation
b1 for TVExpected sales change for extra TV budget, holding radio and newspaper fixed.
b2 for RadioExpected sales change for extra radio budget, holding TV and newspaper fixed.
b3 for NewspaperExpected sales change for extra newspaper budget, holding TV and radio fixed.

Business interpretation

If TV coefficient is high but newspaper coefficient is near zero, what should the marketing team consider?

Takeaway: TV may be more useful, but check data quality, correlation, budget ranges, and business context before making decisions.

TV Radio News weighted sum + b0

Multiple regression is a weighted combination of features plus an intercept.

Why it is still linear

The model is linear in its coefficients. Each coefficient multiplies a feature directly. The features can be many, but the coefficient relationship stays additive.

\[\text{linear: } \beta_0 + \beta_1x_1 + \beta_2x_2\]\[\text{not linear in coefficients: } \beta_0 + \sin(\beta_1x_1)\]

A polynomial feature such as \(x^2\) can still be used inside linear regression:

\[\hat{y} = \beta_0 + \beta_1x + \beta_2x^2\]

This creates a curved relationship in x, but the model is still linear in b0, b1, and b2.

What polynomial regression really does

Polynomial regression does not change the core algorithm. It changes the input features.

Instead of giving the model only \(x\), we give it transformed features:

\[x,\quad x^2,\quad x^3,\quad \ldots\]

Then ordinary linear regression learns one coefficient for each transformed feature.

polynomial fit straight line underfits

Polynomial features let a linear model fit curved relationships.

ModelEquationWhat changes?
Simple linear\(\hat{y}=\beta_0+\beta_1x\)One straight-line feature.
Quadratic\(\hat{y}=\beta_0+\beta_1x+\beta_2x^2\)Adds one curved feature.
Cubic\(\hat{y}=\beta_0+\beta_1x+\beta_2x^2+\beta_3x^3\)Can capture more bends, but risks overfitting.
Polynomial regression can improve underfitting, but high-degree polynomial features can overfit. A curve that passes through every training point may perform badly on new data.

Pattern recognition

Can you think of a relationship that increases first but later slows down?

Examples: experience vs salary, ad spend vs sales, study hours vs marks. These are natural places to consider polynomial features.

Matrix notation of multiple linear regression

For many rows and many features, writing one equation per row becomes messy. Matrix notation lets us write the whole dataset at once.

\[\hat{y} = X\beta\]

Suppose we have 4 training examples and 3 features. The design matrix \(X\) includes a first column of 1s for the intercept.

\[ X = \begin{bmatrix} 1 & x_{11} & x_{12} & x_{13} \\ 1 & x_{21} & x_{22} & x_{23} \\ 1 & x_{31} & x_{32} & x_{33} \\ 1 & x_{41} & x_{42} & x_{43} \end{bmatrix} \]
\[ \beta = \begin{bmatrix} \beta_0 \\ \beta_1 \\ \beta_2 \\ \beta_3 \end{bmatrix} \qquad y = \begin{bmatrix} y_1 \\ y_2 \\ y_3 \\ y_4 \end{bmatrix} \]

Multiplying \(X\beta\) gives all predictions together:

\[ \hat{y} = \begin{bmatrix} 1 & x_{11} & x_{12} & x_{13} \\ 1 & x_{21} & x_{22} & x_{23} \\ 1 & x_{31} & x_{32} & x_{33} \\ 1 & x_{41} & x_{42} & x_{43} \end{bmatrix} \begin{bmatrix} \beta_0 \\ \beta_1 \\ \beta_2 \\ \beta_3 \end{bmatrix} = \begin{bmatrix} \hat{y}_1 \\ \hat{y}_2 \\ \hat{y}_3 \\ \hat{y}_4 \end{bmatrix} \]

The residual vector is:

\[e = y - \hat{y} = y - X\beta\]

OLS chooses the coefficient vector that minimizes squared residuals:

\[\min_{\beta}\|y-X\beta\|^2\]

When the matrix inverse exists, the closed-form OLS solution is:

\[\beta = (X^TX)^{-1}X^Ty\]
Practical intuition: \(X\) contains the data, \(\beta\) contains what the model learns, and \(X\beta\) produces predictions for all rows at once.

Gradient descent with multiple features

In Session 1, gradient descent updated an intercept and one slope. In multiple regression, the same idea updates a full vector of parameters.

\[ \beta = \begin{bmatrix} \beta_0 \\ \beta_1 \\ \beta_2 \\ \vdots \\ \beta_p \end{bmatrix} \qquad \hat{y}=X\beta \]

The cost function measures prediction error across all training rows:

\[ J(\beta)=\frac{1}{n}(y-X\beta)^T(y-X\beta) \]

The gradient tells us how the cost changes if we slightly change each coefficient:

\[ \nabla_{\beta}J(\beta)=\frac{2}{n}X^T(X\beta-y) \]

Then all coefficients move together in the direction that reduces the cost:

\[ \beta := \beta - \alpha\nabla_{\beta}J(\beta) \]

What changes from simple regression?

Instead of asking, “Should the line slope increase or decrease?”, the model asks this for every feature coefficient at once.

If a feature helps reduce error, its coefficient is adjusted. If a feature adds little signal, its coefficient may stay small or become unstable when mixed with similar features.

Prediction surface check

If the learning rate is too large, what will happen to the coefficient path on the cost surface?

The path may jump around the minimum or even move away from it instead of settling smoothly.

coefficient beta1 beta2 minimum start

For many features, the real cost surface has many coefficient directions. This 2D slice shows the core idea.

OLS gives a direct closed-form solution when the matrix setup is manageable. Gradient descent gives an iterative route that becomes especially useful for large datasets, many features, and models beyond ordinary linear regression.
Previous: Metrics Next: Assumptions