如何比较PCA与NMF的预测能力?咨询NMF成分保留方差计算方法
Great question—this is a common pain point when trying to compare PCA and NMF on an apples-to-apples basis, since NMF doesn’t have the neat eigenvalue-based variance metric that PCA does. Let’s break down how you can estimate variance explained for NMF components, and how to align this with your goal of picking components that capture ~95% of the variance:
Approaches to Estimate NMF Component Variance Contribution
1. Reconstruction Error-Based Explained Variance (Most Reliable)
PCA’s explained variance is inherently tied to minimizing reconstruction error, and we can adapt this logic to NMF. Here’s how:
- First, calculate the total squared sum of your original data (this represents the total "signal" we’re trying to capture):
total_squared_sum = np.sum(X ** 2) - For a given number of NMF components
k, compute the reconstruction error (Frobenius norm of the difference between original and reconstructed data):reconstruction_error = ||X - WH||_FwhereWandHare the NMF factor matrices - The overall variance explained by
kcomponents is:explained_variance_ratio = 1 - (reconstruction_error ** 2) / total_squared_sum
To get the variance contribution of individual components, you can:
- Start with 1 component, calculate its explained variance
- Add a second component, subtract the 1-component explained variance from the 2-component value to get the second component’s incremental contribution
- Repeat this process to build a cumulative variance curve, just like PCA’s
2. Orthogonalized Component Contribution (Approximate)
Since PCA’s components are orthogonal, their variance contributions are independent. For NMF, you can approximate this by orthogonalizing the NMF bases first:
- Use a method like Gram-Schmidt to orthogonalize the columns of the NMF basis matrix
W - For each orthogonalized basis vector, calculate its variance contribution using the same logic as PCA:
component_variance = (ortho_W[:,i].T @ X @ X.T @ ortho_W[:,i]) / np.trace(X @ X.T)
Note: This is an approximation—orthogonalizing NMF’s non-negative bases loses the original non-negativity property, so it’s best used only for variance comparison purposes.
3. Normalized Component Magnitude (Quick Rough Estimate)
NMF’s bases are non-negative and often sparse, so the L2 norm of each basis vector can serve as a rough proxy for its importance. You can compute:component_importance = np.linalg.norm(W[:,i]) / np.sum(np.linalg.norm(W, axis=0))
This is fast but less accurate than the reconstruction error method, since it doesn’t directly measure how much variance the component captures in the original data.
Key Implementation Notes
- Unlike PCA, NMF doesn’t have a built-in
explained_variance_ratio_attribute in libraries like scikit-learn, so you’ll need to compute it manually. Here’s a quick code snippet to get you started:
import numpy as np from sklearn.decomposition import NMF from sklearn.preprocessing import MinMaxScaler # NMF requires non-negative input—scale data first if needed scaler = MinMaxScaler() X_scaled = scaler.fit_transform(X) total_squared_sum = np.sum(X_scaled ** 2) target_variance = 0.95 cumulative_variance = 0 k = 0 while cumulative_variance < target_variance: k += 1 nmf = NMF(n_components=k, init='nndsvd', random_state=42) W = nmf.fit_transform(X_scaled) H = nmf.components_ recon_error = nmf.reconstruction_err_ cumulative_variance = 1 - (recon_error ** 2) / total_squared_sum print(f"NMF needs {k} components to explain ~{target_variance*100}% variance")
Final Thoughts
While NMF’s "variance explained" isn’t identical to PCA’s (since NMF optimizes non-negative reconstruction rather than orthogonal variance maximization), the reconstruction error method is the most consistent way to align the two methods for your prediction task comparison. Once you’ve selected the component counts for both that hit the 95% threshold, you can directly compare their performance on your downstream prediction task.
内容的提问来源于stack exchange,提问作者Phil D

