Concept 2

Covariance describes the cloud. Eigenvectors find its axes.

This page builds the mathematical heart of PCA slowly: covariance, covariance matrix, eigenvectors, eigenvalues, and why they solve the PCA problem.

Covariance intuition

Movement question

If study hours increase and marks also increase, are the two features moving together or against each other?

They move together, so covariance is positive.

Covariance measures whether two centered features move together.

\[ \operatorname{Cov}(X,Y)=\frac{1}{n-1}\sum_{i=1}^{n}(x_i-\bar{x})(y_i-\bar{y}) \]
PatternCovarianceMeaning
Both increase togetherPositiveSame direction movement
One increases, other decreasesNegativeOpposite direction movement
No clear linear relationNear zeroNo strong shared direction
positivenegativenear zero

Covariance is about joint movement around the mean.

Covariance matrix

Matrix question

If we have 4 features, how many variance values sit on the diagonal of the covariance matrix?

Four. Each feature has its own variance on the diagonal.

For many features, covariance becomes a matrix. If \(X_c\) is centered data:

\[ \Sigma = \frac{1}{n-1}X_c^TX_c \]

For three features, the matrix looks like this:

\[ \Sigma= \begin{bmatrix} \operatorname{Var}(X_1) & \operatorname{Cov}(X_1,X_2) & \operatorname{Cov}(X_1,X_3)\\ \operatorname{Cov}(X_2,X_1) & \operatorname{Var}(X_2) & \operatorname{Cov}(X_2,X_3)\\ \operatorname{Cov}(X_3,X_1) & \operatorname{Cov}(X_3,X_2) & \operatorname{Var}(X_3) \end{bmatrix} \]

The diagonal stores individual feature spread. The off-diagonal entries store how pairs of features move together.

The covariance matrix is a compact summary of the shape and orientation of the data cloud.

A small covariance calculation

Centered data question

Why is covariance easier to understand after subtracting means?

Because positive and negative deviations show whether features move together around the center.

Suppose two centered features have values:

\[ X_c = [-2,\ 0,\ 2], \qquad Y_c = [-4,\ 0,\ 4] \]

Then:

\[ \operatorname{Cov}(X,Y)=\frac{(-2)(-4)+(0)(0)+(2)(4)}{3-1} =\frac{8+0+8}{2}=8 \]

Positive covariance means the features increase together.

Eigenvectors and eigenvalues

Direction question

If the covariance matrix describes the data cloud, what directions should we extract from it?

The natural axes of the cloud: the longest spread direction first, then the next perpendicular direction.

Eigenvectors and eigenvalues come from this equation:

\[ \Sigma v = \lambda v \]
SymbolMeaning in PCA
\(\Sigma\)Covariance matrix: describes data spread.
\(v\)Eigenvector: a principal component direction.
\(\lambda\)Eigenvalue: amount of variance along that direction.

The equation says: when the covariance matrix acts on \(v\), the direction stays the same. Only its length changes by \(\lambda\).

That is exactly what PCA needs. PCA is searching for stable directions of spread in the covariance matrix. Eigenvectors are those stable directions. Eigenvalues tell how much the data stretches along each stable direction.

\[ \text{eigenvector} = \text{direction of a principal axis} \]
\[ \text{eigenvalue} = \text{variance captured along that axis} \]
large eigenvaluesmall eigenvalue

Eigenvectors give directions. Eigenvalues say how important those directions are.

Why eigenvectors solve PCA

Optimization question

If PCA wants maximum variance, which eigenvector should become PC1?

The eigenvector with the largest eigenvalue.

From the previous page, variance along direction \(w\) is:

\[ \operatorname{Var}(X_cw)=w^T\Sigma w \]

PCA maximizes this while keeping \(w\) as a unit vector:

\[ \max_w\ w^T\Sigma w \quad \text{subject to } w^Tw=1 \]

This optimization leads to:

\[ \Sigma w = \lambda w \]

So the best directions are eigenvectors of the covariance matrix. The best first direction is the eigenvector with the largest eigenvalue.

Intuition without heavy calculus: the covariance matrix already knows the shape of the data cloud. If the cloud is stretched like an ellipse, the longest axis is a direction where the covariance transformation stretches a vector the most. A direction that only gets stretched, without getting rotated away, is an eigenvector. The amount of stretch is the eigenvalue.

PCA needEigen concept that provides it
Find a direction, not an original feature column.Eigenvector gives a new axis.
Rank directions by importance.Eigenvalue gives variance on that axis.
Make components independent/perpendicular.Symmetric covariance matrices produce orthogonal eigenvectors.
Compress by keeping only the best directions.Keep eigenvectors with the largest eigenvalues.

Principal components are perpendicular

Axis question

After PC1 captures the strongest direction, should PC2 repeat the same direction?

No. PC2 must capture a new direction, so it is perpendicular to PC1.

PCA builds a new coordinate system. The component directions are orthogonal:

\[ v_i^Tv_j=0 \quad \text{when } i\neq j \]

Each principal component explains the largest remaining variance after the previous components have already taken their share.

This is why PCA components are uncorrelated with each other, even when the original features were correlated.

Ordering principal components

Ranking question

If one direction explains variance 10 and another explains variance 2, which one should come first?

The direction explaining variance 10.

Sort eigenvalues from largest to smallest:

\[ \lambda_1 \geq \lambda_2 \geq \cdots \geq \lambda_p \]

Then arrange eigenvectors in the same order:

\[ W = [v_1\ v_2\ \cdots\ v_k] \]

\(v_1\) becomes PC1, \(v_2\) becomes PC2, and so on.

Previous: GeometryNext: Projection