Concept 1

PCA begins with geometry: preserve the largest spread.

Before eigenvectors appear, PCA is simply asking which direction gives the best one-dimensional view of the data.

PCA as the best viewing angle

Projection question

If you must flatten a tilted 2D cloud onto one line, should the line go along the cloud or across the cloud?

Along the cloud, because that keeps more spread and loses less information.

PCA searches for a line or axis where projected points remain as spread out as possible. The same line can also be understood as the line that passes closest to the data cloud in squared-distance terms.

The first principal component is the direction with maximum variance:

\[ w_1 = \underset{\|w\|=1}{\arg\max}\ \operatorname{Var}(Xw) \]

Here, \(w\) is not a row from the data. It is a candidate direction in feature space.

ObjectShapeMeaning
\(X\)\(n \times p\)Data matrix with \(n\) rows and \(p\) features.
\(w\)\(p \times 1\)One direction vector, with one weight for each original feature.
\(Xw\)\(n \times 1\)One projected number for every row.

Each value in \(Xw\) is the coordinate or score of one data point along direction \(w\). If those projected scores are widely spread out, then \(w\) is a good direction for preserving information.

\[ Xw= \begin{bmatrix} \text{score of row 1 on direction }w\\ \text{score of row 2 on direction }w\\ \vdots\\ \text{score of row }n\text{ on direction }w \end{bmatrix} \]

The constraint \(\|w\|=1\) prevents a fake solution. Without it, we could make \(Xw\) arbitrarily large just by multiplying \(w\) by a huge number.

large spread small spread

The best 1D summary is the direction where projections are most spread out.

Closest line and closest plane interpretation

Closest subspace question

In 3D data, if we reduce to 2D, what should the chosen plane do?

It should pass as close as possible to the original data cloud.

For 2D to 1D reduction, PCA finds the best line. For 3D to 2D reduction, PCA finds the best plane. In general, PCA finds the best lower-dimensional linear subspace.

Original dimensionReduced dimensionPCA object
2D1DClosest line
3D2DClosest plane
\(p\)D\(k\)DClosest \(k\)-dimensional linear subspace
This interpretation helps students connect PCA with information loss: the farther points are from the chosen line or plane, the more information is lost.
2D to 1D: line3D to 2D: plane

PCA can be seen as finding the closest lower-dimensional flat surface through the centered data.

Why minimum error equals maximum variance

Tradeoff question

If a point is far from the projection line, what happens to reconstruction error?

It increases because the dropped perpendicular part is large.

For a centered point \(x\), split it into the part along the PCA direction and the part perpendicular to it:

\[ x = x_{\parallel} + x_{\perp} \]

Because these two parts are perpendicular, Pythagoras gives:

\[ \|x\|^2 = \|x_{\parallel}\|^2 + \|x_{\perp}\|^2 \]

The total \(\|x\|^2\) is fixed for the dataset. So if we make projected variance \(\|x_{\parallel}\|^2\) large, the leftover reconstruction error \(\|x_{\perp}\|^2\) becomes small.

\[ \text{large projected variance} \Rightarrow \text{small reconstruction error} \]

Step 1: arrange data as a matrix

Matrix question

In a dataset with 5 students and 3 marks columns, what should a row represent and what should a column represent?

A row is one observation. A column is one feature.

We start with a data matrix \(X\):

\[ X = \begin{bmatrix} x_{11} & x_{12} & \cdots & x_{1p}\\ x_{21} & x_{22} & \cdots & x_{2p}\\ \vdots & \vdots & \ddots & \vdots\\ x_{n1} & x_{n2} & \cdots & x_{np} \end{bmatrix} \]

\(n\) is the number of rows or observations. \(p\) is the number of original features.

Step 2: mean centering

Centering question

If we move the entire data cloud to the origin, do distances between points change?

No. The location changes, but the shape and spread remain the same.

PCA studies variation around the center. So we subtract each feature's mean.

\[ \bar{x}_j = \frac{1}{n}\sum_{i=1}^{n}x_{ij} \]
\[ x_{ij}^{centered}=x_{ij}-\bar{x}_j \]

After centering, each feature has mean zero:

\[ \frac{1}{n}\sum_{i=1}^{n}x_{ij}^{centered}=0 \]
before centeringafter centering

Centering moves the average point to the origin. PCA then studies spread around that center.

Step 3: scaling when feature units differ

Scale question

If income is measured in rupees and age is measured in years, which feature may dominate PCA?

Income, because its numeric spread can be much larger.

If features have different units or very different ranges, standardize them before PCA.

\[ z_{ij}=\frac{x_{ij}-\mu_j}{\sigma_j} \]

This makes each feature have mean \(0\) and standard deviation \(1\). PCA then compares patterns more fairly.

unscaled: big unit dominatesscaled: pattern direction appears

Scaling can change the principal components because it changes the shape PCA sees.

What PCA is optimizing

PCA looks for a unit direction \(w\) that maximizes the variance of projected data \(X_cw\):

\[ \operatorname{Var}(X_cw)=\frac{1}{n-1}(X_cw)^T(X_cw) \]

Rearrange this expression:

\[ \operatorname{Var}(X_cw)=w^T\left(\frac{X_c^TX_c}{n-1}\right)w \]

The matrix inside the parentheses is the covariance matrix. That is why covariance is the next major idea.

Back to OverviewNext: Covariance