Principal Component Analysis: reduce dimensions by finding the main directions of variation.
PCA turns many correlated features into fewer new axes that preserve most of the data cloud shape.
Why do we need PCA?
Opening question
If a dataset has 100 features, do all 100 features always add 100 different pieces of information?
Not necessarily. Many features may be correlated, repeated, noisy, or combinations of each other.
Principal Component Analysis is a technique for reducing the number of features while preserving as much variation as possible.
It does this by creating new features called principal components. These components are directions in the data where the data spreads out the most.
A useful mental picture: PCA rotates the coordinate system so that the first axis points along the longest direction of the data cloud.
PC1 captures the largest variation. PC2 captures the next largest variation, perpendicular to PC1.
Two equivalent PCA viewpoints
Viewpoint question
Should PCA be explained as keeping maximum variance or as losing minimum information?
Both. These are two sides of the same idea.
| Viewpoint | Meaning | Classroom intuition |
|---|---|---|
| Maximum variance | Choose directions where projected points are most spread out. | The shadow should still show the shape of the data. |
| Minimum reconstruction error | Choose a lower-dimensional subspace from which original points can be rebuilt as closely as possible. | The projection line or plane should pass close to the cloud. |
This equivalence is one of the most important PCA intuitions. We will revisit it when we discuss projection and reconstruction.
PCA story in one line
Each step has a simple role. Centering places the cloud around the origin. Covariance describes the cloud shape. Eigenvectors find the cloud axes. Projection rewrites each point using those new axes.
Session roadmap
Geometry
Best viewing angle, variance, centering, and scaling.
Math core
Covariance, covariance matrix, eigenvectors, and eigenvalues.
Projection
Transforming data, reconstruction, explained variance, and choosing components.
Practice
Visualization, ML pipelines, noise reduction, use cases, and limitations.
Where PCA is useful
| Use case | How PCA helps |
|---|---|
| Visualization | Convert many features into 2 or 3 components and plot them. |
| Compression | Represent data using fewer numbers while retaining most variation. |
| Noise reduction | Drop very small-variance directions that may mostly contain noise. |
| Multicollinearity | Replace correlated original features with uncorrelated components. |
| Speed | Train later models on fewer dimensions. |
Reference ideas used
This teaching material is strengthened using PCA explanations from Stanford STATS 202 and detailed PCA lecture notes, especially the ideas of closest lower-dimensional subspace, maximum variance, reconstruction error, explained variance, SVD, and whitening.
| Reference | Useful idea |
|---|---|
| Stanford STATS 202 PCA notes | First PC as closest line, PC scores, orthogonal second component, scaled vs unscaled PCA, scree plot. |
| Detailed PCA lecture notes PDF | Projection/reconstruction geometry, variance-error equivalence, projection matrices, dropped-eigenvalue reconstruction error, whitening/sphering. |