Concept 5

Multiple linear regression

Real problems rarely depend on one input. Multiple regression lets many features contribute to one prediction.

Equation

\[\hat{y} = \beta_0 + \beta_1x_1 + \beta_2x_2 + \beta_3x_3 + \cdots + \beta_px_p\]

Every feature gets its own coefficient. The interpretation is:

If x1 increases by one unit while all other features stay constant, the prediction changes by b1 units.

Feature brainstorm

If we are predicting house price, what features should we use?

Group answers into numerical, categorical, useful, suspicious, and unavailable-at-prediction-time features.

Advertising example

\[\widehat{sales} = \beta_0 + \beta_1(TV) + \beta_2(Radio) + \beta_3(Newspaper)\]
CoefficientInterpretation
b1 for TVExpected sales change for extra TV budget, holding radio and newspaper fixed.
b2 for RadioExpected sales change for extra radio budget, holding TV and newspaper fixed.
b3 for NewspaperExpected sales change for extra newspaper budget, holding TV and radio fixed.

Business interpretation

If TV coefficient is high but newspaper coefficient is near zero, what should the marketing team consider?

Expected direction: TV may be more useful, but check data quality, correlation, budget ranges, and business context before making decisions.

TV Radio News weighted sum + b0

Multiple regression is a weighted combination of features plus an intercept.

Why it is still linear

The model is linear in its coefficients. Each coefficient multiplies a feature directly. The features can be many, but the coefficient relationship stays additive.

\[\text{linear: } \beta_0 + \beta_1x_1 + \beta_2x_2\]\[\text{not linear in coefficients: } \beta_0 + \sin(\beta_1x_1)\]

A polynomial feature such as x^2 can still be used inside linear regression:

\[\hat{y} = \beta_0 + \beta_1x + \beta_2x^2\]

This creates a curved relationship in x, but the model is still linear in b0, b1, and b2.

Pattern recognition

Can you think of a relationship that increases first but later slows down?

Examples: experience vs salary, ad spend vs sales, study hours vs marks. Use this to motivate polynomial features.

Matrix form, optional

\[\hat{y} = X\beta\]\[\beta = (X^TX)^{-1}X^Ty\]

Use this only if the batch is comfortable with linear algebra. The practical message is enough: linear regression finds the coefficient vector that minimizes squared error.

Previous: Metrics Next: Assumptions