Concept 3

Projection creates new features. Explained variance tells what we kept.

Once PCA has directions, we project data onto them, measure retained variance, and choose how many components to keep.

Projection onto principal components

Coordinate question

After finding the best new axis, how do we express each old point using that new axis?

We project the point onto the principal component direction.

Let \(W\) contain the top \(k\) principal component directions:

\[ W = [v_1\ v_2\ \cdots\ v_k] \]

The transformed data is:

\[ Z = X_cW \]
ObjectShapeMeaning
\(X_c\)\(n \times p\)Centered original data
\(W\)\(p \times k\)Selected principal component directions
\(Z\)\(n \times k\)New lower-dimensional data

Geometric meaning of projection

Shadow question

If a point casts a shadow onto a line, does the shadow keep all information about the original point?

No. It keeps position along the line and loses distance away from the line.

Projection is like dropping a perpendicular shadow from each point onto a principal component axis.

For a single unit direction \(v\), the projected point in the original coordinate system is:

\[ x_{\parallel}=vv^Tx \]
\[ z_i = x_i^T v_1 \]

For one component, each row becomes one number: the coordinate along PC1.

PC1 linered points are projections

Projection keeps coordinates along the chosen component and drops the perpendicular part.

Scores and loadings

Output question

After PCA, what are the new feature values called, and what tells us how original features combine?

The transformed values are scores. The component directions are loadings.

TermMeaningWhere it appears
LoadingsWeights showing how original features combine to form each component.Columns of \(W\)
ScoresNew coordinates of each row after projection.Rows of \(Z=X_cW\)

If PC1 has high positive loading for two features, then PC1 increases when both features increase together.

\[ \text{PC1 score for row } i = x_{i1}v_{11}+x_{i2}v_{21}+\cdots+x_{ip}v_{p1} \]

Projection matrix view

Subspace question

If we keep two principal components, how do we project a point onto the 2D PCA plane?

Use the matrix containing those two component directions.

Let \(V\) contain the selected orthonormal component directions. Projection into the lower-dimensional coordinates is:

\[ y = V^Tx \]

Reconstruction back into the original feature space is:

\[ \hat{x}=Vy=VV^Tx \]

The matrix \(VV^T\) is the projection matrix onto the subspace spanned by the selected principal components.

This is a powerful geometric idea: PCA is not deleting random coordinates; it is projecting points onto the best lower-dimensional subspace.

Reconstruction after reducing dimensions

Loss question

If we reduce 2D data to 1D and then rebuild it, should we expect the exact original points back?

No. We only get an approximation because one direction was dropped.

After projection, we can approximately reconstruct the centered data:

\[ \hat{X}_c = ZW^T \]

If the original data was centered, add the mean back to reconstruct in the original location:

\[ \hat{X}=\hat{X}_c + \mu \]

The difference between original and reconstructed data is information loss:

\[ \text{Reconstruction Error}=\|X_c-\hat{X}_c\|^2 \]

Reconstruction error from dropped eigenvalues

Dropped variance question

If we keep the first 2 components, where does the lost information come from?

From the components we did not keep.

When eigenvalues are sorted from largest to smallest, keeping the first \(k\) components preserves:

\[ \lambda_1+\lambda_2+\cdots+\lambda_k \]

The leftover reconstruction error is connected to the eigenvalues we drop:

\[ \text{Reconstruction Error} = \sum_{j=k+1}^{p}\lambda_j \]

This gives a clean intuition: the small eigenvalues are the variance we are choosing to discard.

Explained variance ratio

Information question

If PC1 explains 80 percent of the variance, what does that tell us?

One direction captures most of the data's spread.

Each eigenvalue is the variance captured by one principal component. The explained variance ratio is:

\[ \text{Explained Variance Ratio}_k=\frac{\lambda_k}{\sum_{j=1}^{p}\lambda_j} \]

For the first \(m\) components, cumulative explained variance is:

\[ \text{Cumulative Variance}_m=\frac{\sum_{k=1}^{m}\lambda_k}{\sum_{j=1}^{p}\lambda_j} \]
PC1PC2PC3cumulative

The scree plot shows component-wise variance; the curve shows cumulative retained variance.

Detailed explained variance example

Percentage question

If PCA gives eigenvalues 12, 5, 2, and 1, how much variance does PC1 explain?

PC1 explains \(12/(12+5+2+1)=60\%\) of the total variance.

Suppose a dataset has four principal components with eigenvalues:

\[ \lambda_1=12,\quad \lambda_2=5,\quad \lambda_3=2,\quad \lambda_4=1 \]

Total variance is the sum of all eigenvalues:

\[ \text{Total variance}=12+5+2+1=20 \]
ComponentEigenvalueExplained variance ratioCumulative variance
PC112\(12/20=0.60=60\%\)60%
PC25\(5/20=0.25=25\%\)85%
PC32\(2/20=0.10=10\%\)95%
PC41\(1/20=0.05=5\%\)100%

If we keep only PC1 and PC2, we keep:

\[ \frac{12+5}{20}=0.85=85\% \]

If we keep PC1, PC2, and PC3, we keep:

\[ \frac{12+5+2}{20}=0.95=95\% \]

The dropped variance after keeping the first 2 components is:

\[ \lambda_3+\lambda_4=2+1=3 \]

So explained variance helps answer a practical question: how many components can we keep while losing an acceptable amount of information?

Choosing number of components

Choice question

If 2 components explain 92 percent variance, should we always keep only 2?

Not always. It depends on the goal: visualization, compression, or prediction.

MethodHow it worksGood for
Cumulative varianceKeep enough components to explain 90 to 95 percent variance.Compression and general dimensionality reduction.
Scree elbowLook for the point where extra components add little variance.Teaching and exploratory analysis.
Downstream validationTry different component counts and evaluate the final ML model.Prediction tasks.
Visualization needChoose 2 or 3 components.Plotting high-dimensional data.
High explained variance does not guarantee best classification accuracy. PCA is unsupervised and does not look at labels.

Complete PCA algorithm

  1. Start with data matrix \(X\).
  2. Center each feature to get \(X_c\).
  3. Scale features if units or ranges differ.
  4. Compute covariance matrix \(\Sigma=\frac{1}{n-1}X_c^TX_c\).
  5. Find eigenvalues and eigenvectors of \(\Sigma\).
  6. Sort eigenvectors by decreasing eigenvalues.
  7. Select top \(k\) eigenvectors to form \(W\).
  8. Project data: \(Z=X_cW\).
  9. Use \(Z\) for visualization, compression, or modeling.
Previous: CovarianceNext: Practice