PCA技术问询:特征向量求解、A矩阵含义及新数据应用方式
Awesome questions about PCA—let’s break these down clearly so you can apply this confidently:
1. How to Apply PCA to New Data?
PCA is a fit-transform method, meaning you first "train" it on your original dataset, then reuse that trained information to transform new data. Here's the exact workflow:
- First, fit PCA on your training data:
- Calculate the mean of each feature in the training set.
- Compute the covariance matrix of the centered training data (more on this in the second question).
- Solve for eigenvectors and eigenvalues, then select the top k eigenvectors (these are your principal components) based on the largest eigenvalues.
- Next, preprocess and transform new data:
- Center the new data using the training set's mean, not the mean of the new data. This is critical—PCA learns the structure of the training distribution, so you can't shift it with new data's stats.
- Multiply the centered new data by the matrix of top k principal components. This projects the new data onto the same variance-maximizing directions you found from the training set.
For example, in code terms (using Python-like pseudocode):
# After fitting on training data train_mean = training_data.mean(axis=0) top_k_components = selected_eigenvectors # shape: (original_features, k) # Transform new data centered_new_data = new_data - train_mean reduced_new_data = centered_new_data @ top_k_components
2. How to Solve for PCA's Eigenvectors, and Is Matrix A the Raw Data or Covariance Matrix?
First, the straight answer: Matrix A in the eigenvector equation $Av = λv$ is the covariance matrix of the centered raw data, not the raw data matrix itself. Here's why and how to do it:
- PCA's core goal is to find directions where the data has the highest variance. The covariance matrix encodes the variance of each feature and the covariance (linear relationship) between pairs of features—so decomposing this matrix gives us exactly the directions we need.
- Step-by-step to solve for eigenvectors:
- Center the raw data: Subtract the mean of each feature from all samples in the dataset. Let's call this centered matrix $X$.
- Compute the covariance matrix: For a dataset with $m$ samples and $n$ features, the covariance matrix $A$ is calculated as:
$$A = \frac{1}{m-1} X^T X$$
(Using $m-1$ instead of $m$ gives an unbiased estimate of the population covariance; for large $m$, the difference is negligible.) - Perform eigen decomposition: Solve $Av = λv$ to get eigenvalues $λ$ and their corresponding eigenvectors $v$.
- Sort and select: Sort eigenvalues in descending order—their corresponding eigenvectors are the principal components. Pick the top k eigenvectors to use for dimensionality reduction.
A quick side note: In practice, many implementations (like scikit-learn's PCA) use Singular Value Decomposition (SVD) instead of direct eigen decomposition of the covariance matrix. SVD is more numerically stable, especially for high-dimensional data, and the left singular vectors of the centered data matrix $X$ are equivalent to the eigenvectors of the covariance matrix $A$. So you'll often see this used in real-world code, but the underlying principle ties back to the covariance matrix.
内容的提问来源于stack exchange,提问作者PyRsquared

