Concept 4

Using PCA responsibly in real machine learning workflows.

PCA is powerful for visualization, compression, multicollinearity, and noise reduction, but it must be used with care.

PCA for visualization

Visualization question

How can we visualize a dataset with 30 features on a 2D screen?

Use PCA to create two components, then plot PC1 vs PC2.

PCA is commonly used to compress many features into two principal components for plotting.

\[ X_{n\times p} \quad \longrightarrow \quad Z_{n\times 2} \]

If classes separate in the PCA plot, the original feature space likely contains useful structure. If they overlap, classification may be harder or may require nonlinear methods.

PC1PC2

A PCA scatter plot helps inspect high-dimensional structure in 2D.

Biplot intuition

Biplot question

After plotting rows using PC1 and PC2, how can we also understand which original features influenced those axes?

Use loadings. A biplot shows scores and feature directions together.

A PCA scatter plot shows the transformed rows, called scores. A biplot adds arrows for original feature loadings.

Part of biplotMeaning
PointsRows or observations shown using PC1 and PC2 scores.
ArrowsOriginal feature directions based on loadings.
Arrow lengthHow strongly that feature contributes to the plotted components.
Arrow directionWhich side of the PCA plot increases for that feature.
A biplot is excellent for teaching because it connects the new PCA space back to the original columns.

PCA before machine learning models

Pipeline question

Should PCA be fit on the full dataset before train-test split?

No. PCA must be fit only on training data to avoid data leakage.

A correct modeling workflow uses a pipeline:

Split data Scale train Fit PCA on train Train model Evaluate test once
\[ \text{Raw features} \rightarrow \text{Scaler} \rightarrow \text{PCA} \rightarrow \text{ML model} \]

This is especially useful for models affected by dimensionality, such as KNN, logistic regression with many correlated features, or models trained on high-dimensional data.

PCA for multicollinearity

Correlation question

If two features are almost copies of each other, should a model treat them as two independent signals?

No. They mostly repeat the same information.

PCA converts correlated original features into uncorrelated components.

\[ \operatorname{Cov}(Z_i,Z_j)=0 \quad \text{for } i\neq j \]

This can make downstream modeling more stable, though it also makes features less directly interpretable.

PCA for noise reduction

Noise question

If the last few principal components explain tiny variance, what might they contain?

They may contain small details or noise.

When signal is concentrated in high-variance directions, dropping low-variance components can reduce noise.

But this is a judgment call: low variance does not always mean useless.

dropping tiny components can smooth noise

Keeping major components can preserve structure while reducing small noisy variation.

Covariance eigen approach vs SVD

Implementation question

Do libraries always explicitly build the covariance matrix to run PCA?

Not always. Many implementations use Singular Value Decomposition because it is numerically stable.

In teaching, the covariance and eigenvector explanation is the most intuitive path:

\[ \Sigma v=\lambda v \]

In practical implementations, PCA is often computed using SVD on centered data:

\[ X_c = U S V^T \]

The columns of \(V\) give principal component directions, and the singular values in \(S\) are connected to explained variance.

For students, it is enough to know that eigenvectors explain the idea, while SVD is commonly used under the hood for stable computation.

PCA whitening

Whitening question

After PCA rotation, components are uncorrelated. Do they automatically have equal variance?

No. Their variances are the eigenvalues, so whitening rescales them.

PCA first rotates data into principal component coordinates. Whitening goes one step further and divides each component by its standard deviation.

\[ Z = X_cU \]
\[ Z_{white}=Z\Lambda^{-1/2} \]

The goal is to make the transformed data have identity covariance:

\[ \operatorname{Cov}(Z_{white}) \approx I \]

Geometrically, whitening turns an elongated ellipsoid into a more spherical cloud.

PCA scores: unequal variancewhitened: unit variance

Whitening decorrelates and rescales. It can help some algorithms, but may amplify noise.

Whitening is not always better. Dividing by very small eigenvalues can amplify low-variance noise, so it should be used deliberately.

PCA limitations

Limitation question

If the data forms a curved spiral, will straight PCA axes fully capture that shape?

No. PCA is linear, so it captures straight directions of variation.

LimitationWhy it matters
Linear methodCannot directly capture curved structure.
Sensitive to scalingLarge-scale features can dominate components.
Sensitive to outliersOutliers can change variance directions.
UnsupervisedFinds variance, not class separation or target prediction.
Harder interpretabilityComponents are combinations of original features.

How to explain PCA in one minute

PCA looks at the shape of the data cloud. It finds the direction where the cloud is longest, calls that PC1, then finds the next perpendicular direction, calls that PC2, and continues. We can keep only the first few directions if they preserve most of the variation.

\[ \text{Data cloud} \rightarrow \text{main axes} \rightarrow \text{new compact representation} \]
The central intuition: PCA does not choose original columns. It creates new axes by combining original columns.

Final checklist

QuestionWhy it matters
Were features centered?PCA works on variation around the mean.
Were features scaled?Required when feature units differ.
How many components were kept?Controls compression and information loss.
Was PCA fit only on training data?Prevents leakage.
Is PCA suitable for the goal?Variance preservation may not equal predictive usefulness.
Previous: ProjectionBack to Overview