Using PCA responsibly in real machine learning workflows.
PCA is powerful for visualization, compression, multicollinearity, and noise reduction, but it must be used with care.
PCA for visualization
Visualization question
How can we visualize a dataset with 30 features on a 2D screen?
Use PCA to create two components, then plot PC1 vs PC2.
PCA is commonly used to compress many features into two principal components for plotting.
If classes separate in the PCA plot, the original feature space likely contains useful structure. If they overlap, classification may be harder or may require nonlinear methods.
A PCA scatter plot helps inspect high-dimensional structure in 2D.
Biplot intuition
Biplot question
After plotting rows using PC1 and PC2, how can we also understand which original features influenced those axes?
Use loadings. A biplot shows scores and feature directions together.
A PCA scatter plot shows the transformed rows, called scores. A biplot adds arrows for original feature loadings.
| Part of biplot | Meaning |
|---|---|
| Points | Rows or observations shown using PC1 and PC2 scores. |
| Arrows | Original feature directions based on loadings. |
| Arrow length | How strongly that feature contributes to the plotted components. |
| Arrow direction | Which side of the PCA plot increases for that feature. |
PCA before machine learning models
Pipeline question
Should PCA be fit on the full dataset before train-test split?
No. PCA must be fit only on training data to avoid data leakage.
A correct modeling workflow uses a pipeline:
This is especially useful for models affected by dimensionality, such as KNN, logistic regression with many correlated features, or models trained on high-dimensional data.
PCA for multicollinearity
Correlation question
If two features are almost copies of each other, should a model treat them as two independent signals?
No. They mostly repeat the same information.
PCA converts correlated original features into uncorrelated components.
This can make downstream modeling more stable, though it also makes features less directly interpretable.
PCA for noise reduction
Noise question
If the last few principal components explain tiny variance, what might they contain?
They may contain small details or noise.
When signal is concentrated in high-variance directions, dropping low-variance components can reduce noise.
But this is a judgment call: low variance does not always mean useless.
Keeping major components can preserve structure while reducing small noisy variation.
Covariance eigen approach vs SVD
Implementation question
Do libraries always explicitly build the covariance matrix to run PCA?
Not always. Many implementations use Singular Value Decomposition because it is numerically stable.
In teaching, the covariance and eigenvector explanation is the most intuitive path:
In practical implementations, PCA is often computed using SVD on centered data:
The columns of \(V\) give principal component directions, and the singular values in \(S\) are connected to explained variance.
PCA whitening
Whitening question
After PCA rotation, components are uncorrelated. Do they automatically have equal variance?
No. Their variances are the eigenvalues, so whitening rescales them.
PCA first rotates data into principal component coordinates. Whitening goes one step further and divides each component by its standard deviation.
The goal is to make the transformed data have identity covariance:
Geometrically, whitening turns an elongated ellipsoid into a more spherical cloud.
Whitening decorrelates and rescales. It can help some algorithms, but may amplify noise.
PCA limitations
Limitation question
If the data forms a curved spiral, will straight PCA axes fully capture that shape?
No. PCA is linear, so it captures straight directions of variation.
| Limitation | Why it matters |
|---|---|
| Linear method | Cannot directly capture curved structure. |
| Sensitive to scaling | Large-scale features can dominate components. |
| Sensitive to outliers | Outliers can change variance directions. |
| Unsupervised | Finds variance, not class separation or target prediction. |
| Harder interpretability | Components are combinations of original features. |
How to explain PCA in one minute
PCA looks at the shape of the data cloud. It finds the direction where the cloud is longest, calls that PC1, then finds the next perpendicular direction, calls that PC2, and continues. We can keep only the first few directions if they preserve most of the variation.
Final checklist
| Question | Why it matters |
|---|---|
| Were features centered? | PCA works on variation around the mean. |
| Were features scaled? | Required when feature units differ. |
| How many components were kept? | Controls compression and information loss. |
| Was PCA fit only on training data? | Prevents leakage. |
| Is PCA suitable for the goal? | Variance preservation may not equal predictive usefulness. |