使用Numpy处理大量相关随机变量去相关的问题咨询
Hey there! Let's break down your problem and figure out how to get those correlations down to your target level.
First, Why Your Current Approaches Aren't Working
Let's start with the root cause: you've got 1200 variables but only 1000 data points. When the number of variables exceeds the sample size, your sample covariance matrix becomes singular (rank-deficient)—it has at least 200 zero eigenvalues. That's why Cholesky decomposition fails (it requires a strictly positive-definite matrix), and vanilla ZCA whitening struggles to fully eliminate correlations because it can't handle those zero eigenvalues cleanly.
Also, the "unexpectedly high" correlations you're seeing (0.1-0.2) aren't a fluke. In high-dimensional settings where variables outnumber samples, sample covariance estimates are inherently noisy—even if your true data is independent, sample correlations will have non-trivial biases.
A Practical Solution: Regularized SVD-Based Whitening
The best approach here is to use SVD (Singular Value Decomposition) instead of Cholesky, since SVD works seamlessly with singular matrices. Adding a small regularization term will stabilize the inverse and push off-diagonal correlations to near-zero. Here's a step-by-step implementation:
import numpy as np # Your original data generation code seed = 0 sample_size = 1000 n_var = 1200 total_rng = np.random.RandomState(seed=seed).randn(sample_size*n_var).reshape((n_var, sample_size)) # Step 1: Center the data (remove variable means) centered_data = total_rng - total_rng.mean(axis=1, keepdims=True) # Step 2: Perform SVD on the centered data U, S, Vt = np.linalg.svd(centered_data, full_matrices=False) # Step 3: Add a tiny regularization epsilon to avoid division by zero # Adjust eps if needed—1e-8 is a safe starting point eps = 1e-8 # Calculate the inverse square root of the scaled singular values # This accounts for the sample covariance scaling factor (sample_size - 1) inv_sqrt_scaled_S = 1.0 / np.sqrt((S ** 2) / (sample_size - 1) + eps) # Step 4: Build the regularized ZCA whitening matrix whitening_matrix = U @ np.diag(inv_sqrt_scaled_S) @ U.T # Step 5: Apply the whitening transformation whitened_data = whitening_matrix @ centered_data # Verify the result: check max off-diagonal correlation cov_matrix = np.cov(whitened_data) max_off_diag_corr = np.max(np.abs(cov_matrix - np.eye(n_var))) print(f"Max off-diagonal correlation after whitening: {max_off_diag_corr:.6f}")
What This Does
- Centering: Ensures each variable has a mean of 0, so correlations match covariance values.
- SVD: Decomposes your data into orthogonal components, which lets us directly manipulate the variance/covariance structure.
- Regularization: The
epsterm prevents division by zero for near-zero singular values (which come from the rank deficiency) and adds stability to the transformation. - ZCA Whitening: Unlike PCA whitening (which rotates data to uncorrelated axes), ZCA keeps the data in the original variable space while making variables uncorrelated (covariance matrix close to identity).
Can We Get Correlations Exactly to 0?
In short: no, not with 1200 variables and only 1000 samples. The mathematical reason is that your data lives in a 1000-dimensional space (since you have 1000 data points), so you can't have 1200 linearly independent variables—200 of them will always be linear combinations of the others. That means tiny residual correlations might remain, but with the regularized SVD approach, you can easily get them below 1e-4 (the code above should get you there with the right eps).
Quick Tips to Tune the Result
- If the max correlation is still too high, try increasing
epsslightly (e.g., 1e-7) to add more regularization. - Double-check that you're computing correlations correctly—remember that covariance equals correlation for mean-zero, unit-variance variables (which whitening achieves).
内容的提问来源于stack exchange,提问作者bombdruid

