如何缩减10000×10000相关矩阵并提取阈值≥0.6的相关值
Hey there! Let's break down how to extract those correlation values above 0.6 from your large matrix, building on the code you’ve already written. First, let’s recap what your existing code does, then dive into extracting the specific high-correlation pairs and values.
What Your Current Code Achieves
Your code already does the heavy lifting of reducing the massive 10000×10000 matrix to a smaller subset:
- You generate your data
cellA, compute the transposed correlation matrixcorr - You set the diagonal (self-correlations) to 0, then filter for genes that have at least one correlation with absolute value ≥ 0.6
- The result is
filter.corr.gene—a trimmed-down correlation matrix containing only those relevant genes. This is perfect for generating a manageable heatmap (trypheatmap::pheatmap(filter.corr.gene)for a polished version!)
Extracting the Specific High-Correlation Values
If you want to pull out the actual pairs of genes and their corresponding correlation values that meet the ≥0.6 threshold, here are two straightforward methods:
Method 1: Using which() to Locate High-Correlation Positions
This approach directly finds the indices of values that meet your threshold and converts them into a readable data frame:
# Define your threshold (you already have this) threshold <- 0.6 # Get row/column indices of values ≥ threshold (absolute value) in the filtered matrix high_corr_indices <- which(abs(filter.corr.gene) >= threshold, arr.ind = TRUE) # Extract the actual correlation values high_corr_values <- filter.corr.gene[high_corr_indices] # Combine into a tidy data frame high_corr_results <- data.frame( Gene1 = rownames(filter.corr.gene)[high_corr_indices[, "row"]], Gene2 = colnames(filter.corr.gene)[high_corr_indices[, "col"]], Correlation = high_corr_values ) # Optional: Remove duplicate symmetric pairs (e.g., Gene1-Gene2 and Gene2-Gene1) high_corr_results <- high_corr_results[high_corr_results$Gene1 < high_corr_results$Gene2, ]
Method 2: Reshape the Matrix to Long Format with reshape2
If you prefer a more "tidy data" approach, melting the matrix into long format makes filtering easy:
# Install and load reshape2 if you haven't already # install.packages("reshape2") library(reshape2) # Convert the filtered correlation matrix to long format melted_corr <- melt(filter.corr.gene) colnames(melted_corr) <- c("Gene1", "Gene2", "Correlation") # Filter for values ≥ threshold (exclude self-correlations, though you already set diag to 0) high_corr_results <- subset(melted_corr, abs(Correlation) >= threshold & Gene1 != Gene2) # Optional: Remove symmetric duplicates high_corr_results <- high_corr_results[high_corr_results$Gene1 < high_corr_results$Gene2, ]
Quick Notes
- If you only care about positive correlations (not absolute values), replace
abs(Correlation) >= thresholdwithCorrelation >= threshold. For negative correlations, useCorrelation <= -threshold. - Your initial step of filtering down to
filter.corr.geneis crucial—processing the full 10000×10000 matrix directly would be slow and memory-intensive, so this pre-filter saves a ton of resources.
内容的提问来源于stack exchange,提问作者adR

