寻找对两向量相关性有贡献的子空间及时间序列相关子区间分析
Got it, let's tackle your two problems head-on—both focus on pinpointing subspaces where correlation actually matters, even when the overall data looks uncorrelated. Here's a practical, step-by-step breakdown for each:
If you're working with two high-dimensional vectors (or sets of vectors), the key is to isolate the orthogonal subspaces where their covariance is strongest. Here's the standard workflow:
- Start with covariance decomposition: Compute the covariance matrix $\Sigma = \text{Cov}(X, Y)$ between your two vector sets. Then perform a Singular Value Decomposition (SVD) on $\Sigma$, which breaks it down into $\Sigma = U\Sigma V^T$.
- The diagonal elements of $\Sigma$ are the singular values—these quantify the amount of covariance explained by each corresponding pair of subspaces (from $U$ and $V$).
- Identify significant subspaces: Rank the singular values from largest to smallest. Subspaces associated with the top k singular values are the ones driving the correlation. To decide k:
- Use a variance cutoff: Pick k such that the cumulative sum of singular values accounts for, say, 95% of the total variance (adjust based on your needs).
- Use statistical testing: Compare each singular value to the distribution of singular values you'd get from random, uncorrelated vectors. If a singular value is significantly larger than the 95th percentile of the random distribution, it's a meaningful subspace.
For time series that only correlate in narrow windows (like midnight minutes), you need a combination of sliding window analysis and statistical validation to avoid mistaking random noise for real correlation:
Step 1: Pinpoint candidate time intervals
- Use a sliding window approach: Define a window size matching your expected correlated interval (e.g., 5-minute windows for your midnight example). Slide this window across your entire time series, calculating a correlation metric (like Pearson's r) for the $x_i$ and $y_i$ in each window.
- Flag outliers: Identify windows where the correlation coefficient is far outside the range of coefficients you get from the rest of the uncorrelated data (e.g., 2 standard deviations above the mean of all window correlations).
Step 2: Calculate P-values for candidate windows
For each flagged window, test if the correlation is statistically significant (not just random):
- Parametric test (for normal data): Compute the t-statistic from Pearson's r:
where n is the number of samples in the window. Then use a t-distribution with (n-2) degrees of freedom to find the two-tailed P-value (since correlation could be positive or negative).t = r * sqrt((n - 2) / (1 - r²)) - Non-parametric permutation test (robust to non-normal data):
- Shuffle the values of one time series (e.g., $y_i$) to break any temporal correlation.
- Recalculate the correlation for the candidate window with the shuffled data.
- Repeat this 1000+ times to build a distribution of random correlation coefficients.
- The P-value is the fraction of permuted correlations that are as extreme or more extreme than your original window's correlation.
Step 3: Account for multiple testing
Since you're testing dozens/hundreds of sliding windows, you're at risk of false positives. Fix this with:
- Bonferroni correction: Multiply each P-value by the total number of windows tested (conservative, but reduces false positives).
- FDR correction: Use methods like Benjamini-Hochberg to adjust P-values while balancing false positives and true discoveries.
内容的提问来源于stack exchange,提问作者mox

