Concept 3

OLS and gradient descent

Both stories answer the same question: which slope and intercept produce the smallest overall error?

Ordinary Least Squares

OLS means choose the coefficients that minimize the sum of squared residuals.

\[\min_{\beta_0,\beta_1}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\]

For simple linear regression, the closed-form slope can be written as:

\[\beta_1 = \frac{\sum_{i=1}^{n}(x_i - \bar{x})(y_i - \bar{y})}{\sum_{i=1}^{n}(x_i - \bar{x})^2}\]\[\beta_0 = \bar{y} - \beta_1\bar{x}\]

Intuition: slope is based on how x and y move together, divided by how much x itself varies.

Check your intuition

If x and y usually increase together, should the slope be positive or negative?

What if x increases and y usually decreases?

Cost function intuition: move downhill until you reach the lowest error.

OLS slope intuition

The slope formula can be read as a ratio:

\[\beta_1 = \frac{\text{how }x\text{ and }y\text{ move together}}{\text{how much }x\text{ moves by itself}}\]
\[\beta_1 = \frac{\mathrm{covariance\ numerator}}{\mathrm{variance\ numerator}}\]

Numerator: movement together

\[\sum_{i=1}^{n}(x_i - \bar{x})(y_i - \bar{y})\]

If a point is above average in x and also above average in y, the product is positive. If both are below average, the product is also positive. This pushes the slope upward.

If x is above average but y is below average, the product is negative. This pushes the slope downward.

mean x mean y + product + product - product - product

OLS slope is positive when the positive products dominate the negative products.

The intercept then positions the line so it passes through the average point \((\bar{x}, \bar{y})\): \(\beta_0 = \bar{y} - \beta_1\bar{x}\).

Gradient descent

Gradient descent does not jump directly to the answer. It starts with guesses for b0 and b1, checks the error, and repeatedly updates the guesses.

\[\theta_{new} = \theta_{old} - \alpha\nabla J(\theta)\]
Guess b0, b1 Predict y_hat Calculate MSE Find gradient Update values

Gradient derivation for simple linear regression

Start with the model equation:

\[\hat{y}_i = \beta_0 + \beta_1x_i\]

Use Mean Squared Error as the cost function:

\[J(\beta_0,\beta_1) = \frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2\]

Substitute the prediction equation into the cost function:

\[J(\beta_0,\beta_1) = \frac{1}{n}\sum_{i=1}^{n}(y_i - (\beta_0 + \beta_1x_i))^2\]

The gradient tells us how the cost changes when we slightly change the intercept or slope.

\[\frac{\partial J}{\partial \beta_0} = -\frac{2}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)\]
\[\frac{\partial J}{\partial \beta_1} = -\frac{2}{n}\sum_{i=1}^{n}x_i(y_i - \hat{y}_i)\]

So the update rules become:

\[\beta_0 := \beta_0 - \alpha\frac{\partial J}{\partial \beta_0}\]
\[\beta_1 := \beta_1 - \alpha\frac{\partial J}{\partial \beta_1}\]
Intuition: if the derivative is positive, increasing the parameter increases error, so gradient descent moves the parameter down. If the derivative is negative, increasing the parameter reduces error, so gradient descent moves the parameter up.
TermMeaningImportant note
GradientDirection and steepness of change in cost.It tells us how to move the parameter to reduce error.
Learning rateStep size used during each update.Too small is slow. Too large can overshoot.
EpochOne full pass over the training data.More epochs may help, but after convergence they add little value.
ConvergenceCost stops improving meaningfully.The model has reached a stable minimum.

Gradient descent intuition

Imagine you are on a hill and can only feel the slope under your feet. How do you reach the lowest point?

This maps directly to gradient direction, step size, overshooting, and convergence.

Interactive gradient descent

Change the initial line and learning settings. The chart shows how gradient descent updates the line and how the cost changes over epochs.

Fitted line

Cost over epochs

Current slope0.00
Current intercept0.00
Final MSE0.00

Convex vs non-convex cost

Linear regression with MSE has a convex cost surface: one smooth bowl and one global minimum. More complex models can have non-convex surfaces with many dips.

Linear regression cost

global minimum

Convex: gradient descent has a clear destination.

Complex model cost

local global local

Non-convex: the starting point can affect where optimization ends.

Important distinction

Scikit-learn's standard LinearRegression uses a direct least-squares solver, not ordinary gradient descent. Still teach gradient descent because it becomes essential for neural networks and many large-scale optimization problems.
Previous: Errors Next: Session 2