PCA begins with geometry: preserve the largest spread.
Before eigenvectors appear, PCA is simply asking which direction gives the best one-dimensional view of the data.
PCA as the best viewing angle
Projection question
If you must flatten a tilted 2D cloud onto one line, should the line go along the cloud or across the cloud?
Along the cloud, because that keeps more spread and loses less information.
PCA searches for a line or axis where projected points remain as spread out as possible. The same line can also be understood as the line that passes closest to the data cloud in squared-distance terms.
The first principal component is the direction with maximum variance:
Here, \(w\) is not a row from the data. It is a candidate direction in feature space.
| Object | Shape | Meaning |
|---|---|---|
| \(X\) | \(n \times p\) | Data matrix with \(n\) rows and \(p\) features. |
| \(w\) | \(p \times 1\) | One direction vector, with one weight for each original feature. |
| \(Xw\) | \(n \times 1\) | One projected number for every row. |
Each value in \(Xw\) is the coordinate or score of one data point along direction \(w\). If those projected scores are widely spread out, then \(w\) is a good direction for preserving information.
The constraint \(\|w\|=1\) prevents a fake solution. Without it, we could make \(Xw\) arbitrarily large just by multiplying \(w\) by a huge number.
The best 1D summary is the direction where projections are most spread out.
Closest line and closest plane interpretation
Closest subspace question
In 3D data, if we reduce to 2D, what should the chosen plane do?
It should pass as close as possible to the original data cloud.
For 2D to 1D reduction, PCA finds the best line. For 3D to 2D reduction, PCA finds the best plane. In general, PCA finds the best lower-dimensional linear subspace.
| Original dimension | Reduced dimension | PCA object |
|---|---|---|
| 2D | 1D | Closest line |
| 3D | 2D | Closest plane |
| \(p\)D | \(k\)D | Closest \(k\)-dimensional linear subspace |
PCA can be seen as finding the closest lower-dimensional flat surface through the centered data.
Why minimum error equals maximum variance
Tradeoff question
If a point is far from the projection line, what happens to reconstruction error?
It increases because the dropped perpendicular part is large.
For a centered point \(x\), split it into the part along the PCA direction and the part perpendicular to it:
Because these two parts are perpendicular, Pythagoras gives:
The total \(\|x\|^2\) is fixed for the dataset. So if we make projected variance \(\|x_{\parallel}\|^2\) large, the leftover reconstruction error \(\|x_{\perp}\|^2\) becomes small.
Step 1: arrange data as a matrix
Matrix question
In a dataset with 5 students and 3 marks columns, what should a row represent and what should a column represent?
A row is one observation. A column is one feature.
We start with a data matrix \(X\):
\(n\) is the number of rows or observations. \(p\) is the number of original features.
Step 2: mean centering
Centering question
If we move the entire data cloud to the origin, do distances between points change?
No. The location changes, but the shape and spread remain the same.
PCA studies variation around the center. So we subtract each feature's mean.
After centering, each feature has mean zero:
Centering moves the average point to the origin. PCA then studies spread around that center.
Step 3: scaling when feature units differ
Scale question
If income is measured in rupees and age is measured in years, which feature may dominate PCA?
Income, because its numeric spread can be much larger.
If features have different units or very different ranges, standardize them before PCA.
This makes each feature have mean \(0\) and standard deviation \(1\). PCA then compares patterns more fairly.
Scaling can change the principal components because it changes the shape PCA sees.
What PCA is optimizing
PCA looks for a unit direction \(w\) that maximizes the variance of projected data \(X_cw\):
Rearrange this expression:
The matrix inside the parentheses is the covariance matrix. That is why covariance is the next major idea.