Multiple linear regression
Real problems rarely depend on one input. Multiple regression lets many features contribute to one prediction.
Equation
Every feature gets its own coefficient. The interpretation is:
Feature brainstorm
If we are predicting house price, what features should we use?
Think about numerical features, categorical features, useful features, suspicious features, and features that may not be available at prediction time.
How does the multivariate model fit?
The model starts with many possible planes or hyperplanes. Each choice of coefficients creates a different prediction surface and different residuals.
The best fit is the surface that minimizes the total squared residuals across all rows:
Intercept moves the surface
\(\beta_0\) shifts the whole line, plane, or hyperplane up and down without changing its tilt.
Coefficients tilt the surface
Each \(\beta_j\) controls tilt in one feature direction. Larger magnitude means the prediction changes more along that feature.
All features compete together
The model chooses coefficients jointly, not one by one. A feature's coefficient depends on the other features present.
Fitting means adjusting the surface until the residual distances are as small as possible overall.
Coefficient meaning: holding other features constant
In multiple regression, each coefficient is a partial effect. It answers:
This is why coefficient interpretation is more subtle than simple regression.
| Question | Simple regression | Multiple regression |
|---|---|---|
| What does slope mean? | How \(y\) changes as one \(x\) changes. | How \(y\) changes as one \(x_j\) changes while other features stay constant. |
| Can coefficients change when we add features? | No other features exist. | Yes. Adding/removing features can change coefficients. |
| Why? | One feature explains everything it can. | Features share explanatory responsibility. |
Projection intuition
Another way to think about fitting is projection. The target vector \(y\) may not be perfectly represented by the feature columns in \(X\). Linear regression finds the closest possible prediction vector \(\hat{y}\) inside the space created by those features.
At the OLS solution, the residual vector is orthogonal to the feature space:
The prediction is the closest point the linear model can reach using the available features.
Why this helps
If important features are missing, the feature space cannot reach the true target well. The residual remains large.
If we add useful features, the feature space expands, and the model may get closer to \(y\).
If we add noisy or redundant features, the model may become harder to interpret or overfit.
From line fitting to plane fitting
With one input feature, linear regression fits a line:
With two input features, it fits a plane:
Example: house price may depend on both area and number of bedrooms.
The model is still linear because the coefficients are added together in a straight weighted sum.
For two features, predictions lie on a plane. Residuals are vertical distances from points to the plane.
Advertising example
| Coefficient | Interpretation |
|---|---|
| b1 for TV | Expected sales change for extra TV budget, holding radio and newspaper fixed. |
| b2 for Radio | Expected sales change for extra radio budget, holding TV and newspaper fixed. |
| b3 for Newspaper | Expected sales change for extra newspaper budget, holding TV and radio fixed. |
Business interpretation
If TV coefficient is high but newspaper coefficient is near zero, what should the marketing team consider?
Takeaway: TV may be more useful, but check data quality, correlation, budget ranges, and business context before making decisions.
Multiple regression is a weighted combination of features plus an intercept.
Why it is still linear
The model is linear in its coefficients. Each coefficient multiplies a feature directly. The features can be many, but the coefficient relationship stays additive.
A polynomial feature such as \(x^2\) can still be used inside linear regression:
This creates a curved relationship in x, but the model is still linear in b0, b1, and b2.
What polynomial regression really does
Polynomial regression does not change the core algorithm. It changes the input features.
Instead of giving the model only \(x\), we give it transformed features:
Then ordinary linear regression learns one coefficient for each transformed feature.
Polynomial features let a linear model fit curved relationships.
| Model | Equation | What changes? |
|---|---|---|
| Simple linear | \(\hat{y}=\beta_0+\beta_1x\) | One straight-line feature. |
| Quadratic | \(\hat{y}=\beta_0+\beta_1x+\beta_2x^2\) | Adds one curved feature. |
| Cubic | \(\hat{y}=\beta_0+\beta_1x+\beta_2x^2+\beta_3x^3\) | Can capture more bends, but risks overfitting. |
Pattern recognition
Can you think of a relationship that increases first but later slows down?
Examples: experience vs salary, ad spend vs sales, study hours vs marks. These are natural places to consider polynomial features.
Matrix notation of multiple linear regression
For many rows and many features, writing one equation per row becomes messy. Matrix notation lets us write the whole dataset at once.
Suppose we have 4 training examples and 3 features. The design matrix \(X\) includes a first column of 1s for the intercept.
Multiplying \(X\beta\) gives all predictions together:
The residual vector is:
OLS chooses the coefficient vector that minimizes squared residuals:
When the matrix inverse exists, the closed-form OLS solution is:
Gradient descent with multiple features
In Session 1, gradient descent updated an intercept and one slope. In multiple regression, the same idea updates a full vector of parameters.
The cost function measures prediction error across all training rows:
The gradient tells us how the cost changes if we slightly change each coefficient:
Then all coefficients move together in the direction that reduces the cost:
What changes from simple regression?
Instead of asking, “Should the line slope increase or decrease?”, the model asks this for every feature coefficient at once.
If a feature helps reduce error, its coefficient is adjusted. If a feature adds little signal, its coefficient may stay small or become unstable when mixed with similar features.
Prediction surface check
If the learning rate is too large, what will happen to the coefficient path on the cost surface?
The path may jump around the minimum or even move away from it instead of settling smoothly.
For many features, the real cost surface has many coefficient directions. This 2D slice shows the core idea.