关于《The Elements of Statistical Learning》中PCA第一主成分方差含1/N的推导疑问
Hey there! Don’t worry—this is totally a common confusion with notation differences across stats resources, so no silly question here at all. Let’s walk through why that 1/N shows up in ESL’s PCA variance calculation.
First, let’s anchor ourselves to the notation in The Elements of Statistical Learning (ESL):
- The data matrix $X$ is $N \times p$, where each row is one of $N$ observations, and each column is a $p$-dimensional feature. Crucially, ESL assumes $X$ is centered (mean of each feature is 0) here.
- When they talk about the variance of the first principal component, they’re referring to the sample variance of the projected values.
Let’s break down the math step by step:
- The projection of the data onto a unit vector $v$ (the principal component direction) is the vector $Xv$. Each entry in this vector is the projection of one observation onto $v$.
- For centered data, the sample variance of these projections is calculated as:
$$
\text{Var}(Xv) = \frac{1}{N} \sum_{i=1}^N (x_i^T v - \mu)^2
$$
Since the data is centered, $\mu = 0$, so this simplifies to:
$$
\text{Var}(Xv) = \frac{1}{N} \sum_{i=1}^N (x_i^T v)^2 = \frac{1}{N} (Xv)^T (Xv) = \frac{1}{N} v^T X^T X v
$$ - Now substitute the SVD of $X$: $X = UDV^T$. Remember that $U$ and $V$ are orthogonal matrices ($U^T U = I$, $V^T V = I$), so $X^T X = V D^2 V^T$. Plugging this in:
$$
\text{Var}(Xv) = \frac{1}{N} v^T V D^2 V^T v
$$ - To maximize this variance (the goal of PCA), we choose $v$ to be the first column of $V$ (the eigenvector corresponding to the largest eigenvalue of $X^T X$). Let’s call this vector $v_1$. Then $V^T v_1 = e_1$ (the standard basis vector with 1 in the first position, 0 elsewhere). Substituting this in:
$$
\text{Var}(Xv_1) = \frac{1}{N} e_1^T D^2 e_1 = \frac{1}{N} d_1^2
$$
There’s that 1/N factor!
The key reason you might not have seen this in other PCA resources is that some textbooks use the unbiased sample covariance (which divides by $N-1$ instead of $N$). But ESL opts for the $N$-scaled version here—this is a deliberate choice in their notation for the sample covariance matrix (they define it as $\frac{1}{N} X^T X$ for centered $X$), which directly leads to the 1/N in the principal component variance.
If you verify with the individual feature variances, you’ll see consistency: the sample variance of a single centered feature is $\frac{1}{N} \sum_{i=1}^N x_{ij}^2$, which is exactly the $j$-th diagonal entry of $\frac{1}{N} X^T X$. The principal component variances are just the eigenvalues of this scaled covariance matrix, hence the 1/N.
备注:内容来源于stack exchange,提问作者chenqile

