OLS and gradient descent
Both stories answer the same question: which slope and intercept produce the smallest overall error?
Ordinary Least Squares
OLS means choose the coefficients that minimize the sum of squared residuals.
For simple linear regression, the closed-form slope can be written as:
Intuition: slope is based on how x and y move together, divided by how much x itself varies.
Check your intuition
If x and y usually increase together, should the slope be positive or negative?
What if x increases and y usually decreases?
Cost function intuition: move downhill until you reach the lowest error.
OLS slope intuition
The slope formula can be read as a ratio:
Numerator: movement together
If a point is above average in x and also above average in y, the product is positive. If both are below average, the product is also positive. This pushes the slope upward.
If x is above average but y is below average, the product is negative. This pushes the slope downward.
OLS slope is positive when the positive products dominate the negative products.
Gradient descent
Gradient descent does not jump directly to the answer. It starts with guesses for b0 and b1, checks the error, and repeatedly updates the guesses.
Gradient derivation for simple linear regression
Start with the model equation:
Use Mean Squared Error as the cost function:
Substitute the prediction equation into the cost function:
The gradient tells us how the cost changes when we slightly change the intercept or slope.
So the update rules become:
| Term | Meaning | Important note |
|---|---|---|
| Gradient | Direction and steepness of change in cost. | It tells us how to move the parameter to reduce error. |
| Learning rate | Step size used during each update. | Too small is slow. Too large can overshoot. |
| Epoch | One full pass over the training data. | More epochs may help, but after convergence they add little value. |
| Convergence | Cost stops improving meaningfully. | The model has reached a stable minimum. |
Gradient descent intuition
Imagine you are on a hill and can only feel the slope under your feet. How do you reach the lowest point?
This maps directly to gradient direction, step size, overshooting, and convergence.
Interactive gradient descent
Change the initial line and learning settings. The chart shows how gradient descent updates the line and how the cost changes over epochs.
Fitted line
Cost over epochs
Convex vs non-convex cost
Linear regression with MSE has a convex cost surface: one smooth bowl and one global minimum. More complex models can have non-convex surfaces with many dips.
Linear regression cost
Convex: gradient descent has a clear destination.
Complex model cost
Non-convex: the starting point can affect where optimization ends.