Learn · Intermediate
Principal component analysis: finding the directions that summarize your data
Principal component analysis finds new perpendicular directions through a dataset, ordered by how much variation each direction captures. Keeping only the leading directions compresses the data into fewer coordinates, which makes patterns easier to inspect and can reduce storage or modeling work. The method preserves variation, not necessarily meaning, predictive usefulness or causal structure.
Turning correlated measurements into coordinates
Imagine measuring the height and arm span of many people. Taller people usually have longer arms, so the plotted points form an elongated cloud rather than a circle. The horizontal and vertical axes treat height and arm span separately. A more informative first axis runs along the cloud’s long direction, summarizing their shared increase; a second axis crosses it and captures departures from the usual relationship.
Principal component analysis, usually shortened to PCA, finds those axes systematically. The axes are linear combinations of the original measurements. A person’s new coordinates are obtained by projecting their measurements onto them. Keeping both axes just changes the coordinate system. Keeping only the first discards the variation along the second and gives a compressed approximation.
Karl Pearson’s 1901 paper on lines and planes of closest fit introduced the geometric approximation perspective. John Shlens’s A Tutorial on Principal Component Analysis gives a modern account connecting geometry, covariance and linear algebra. These sources describe a foundational statistical method rather than a neural network that learns a complicated nonlinear code.
What the calculation actually does
Begin with a table whose rows are examples and columns are measurements. Subtract each column’s mean so the cloud is centered at the origin. Then find the direction along which the centered data have the greatest variance. That is the first principal component. Find the next highest-variance direction subject to being perpendicular to the first, and continue.
One calculation uses the covariance matrix, which records how measurements vary together. Its eigenvectors supply the component directions, and its eigenvalues describe the variation along them. Another uses singular value decomposition directly on the centered data matrix. These are related computational routes to the same ordinary PCA solution, with numerical and efficiency considerations affecting the implementation choice.
Projection produces a score for each example on each retained component. Reconstruction takes those scores, moves back along the retained directions and adds the mean. For a chosen number of components, ordinary PCA minimizes squared reconstruction error among linear subspaces of that size. That statement specifies both its strength and its limits: linear approximation and squared error are the objective.
Why units and preprocessing change the result
Suppose a dataset contains income in dollars and age in years. The numerical spread of income can dominate even if age matters more to the question. Converting dollars to cents changes that spread again. PCA applied directly to covariance is therefore sensitive to measurement scale.
Standardizing columns to comparable variance is a common choice, but it is a modeling decision. It gives a low-variance measurement more influence and can amplify noise. Centering and scaling solve different problems: centering changes the origin, while scaling changes the relative importance of measurements. Neither should be chosen without considering what the data represent.
Choosing what to keep
The explained-variance ratio tells you how much of the dataset’s total variation a component captures. If two components account for most variation, a two-dimensional plot may show the cloud reasonably well. It does not follow that those components preserve all useful information.
Consider a production sensor whose rare failures alter one small measurement while ordinary temperature variation dominates everything else. PCA may discard the failure signal because it contributes little total variance. A classifier can care intensely about a direction that PCA ranks low. Unsupervised compression answers a different question from predicting a label, as the discriminative-model lesson explains.
Select components using the downstream objective as well as reconstruction. Fit the mean, scaling and directions on training data, then apply that same transformation to held-out examples. Fitting them on the full dataset leaks information from evaluation into preprocessing. The holdout-set lesson explains why a seemingly harmless transformation still belongs inside the training procedure.
What a direction does—and does not—explain
PCA is useful for inspecting embeddings and other high-dimensional representations. Yet a leading component does not automatically correspond to an understandable feature. High variance can reflect document length, formatting, measurement noise or another incidental factor. Components with equal or nearly equal variance can also have unstable individual orientations; a direction’s sign is arbitrary.
Today’s NEEDLE paper and its code repository discuss backdoor-related directions and weight editing. That connection illustrates why directional analysis is useful, but NEEDLE is not simply PCA. Finding variation, finding a direction associated with a trigger and demonstrating that an intervention changes behavior are separate tasks.
Goodfire’s geometry discussion offers a related reason to avoid treating every direction as a named concept: representations can have curved, multidimensional structure. A plotted cluster or principal direction is descriptive evidence. A causal claim requires interventions or additional assumptions. PCA gives a compact map of where the data vary; it does not tell you why that variation exists. Used with clear preprocessing, held-out checks and a defined downstream purpose, it is a rigorous and inexpensive first view of a complicated dataset.
Karl Pearson: On lines and planes of closest fit to systems of points in space (1901)
John Shlens: A Tutorial on Principal Component Analysis (2014)
Key questions
What does the first principal component maximize?
Does PCA automatically keep the information a classifier needs?
Should PCA be fitted before or after splitting training and test data?
Cite this
APA
Ground Truth. (2026, October 5). Principal component analysis: finding the directions that summarize your data. Ground Truth. https://groundtruth.day/learn/principal-component-analysis.html
BibTeX
@misc{groundtruth:principal-component-analysis,
title = {Principal component analysis: finding the directions that summarize your data},
author = {{Ground Truth}},
year = {2026},
month = {oct},
url = {https://groundtruth.day/learn/principal-component-analysis.html}
}