如何从heatmaply层级聚类第1阶段簇提取测试样本t1的相似样本
Hey there! Let's walk through exactly how to extract those similar samples to t1 and their features using hclust, dist, and a deep dive into the merge property you're confused about. I’ll keep this practical with code examples so you can follow along.
Step 1: Align Your Clustering with heatmaply
First, critical note: Make sure your manual clustering matches what heatmaply used. Heatmaply defaults to:
- Distance metric:
euclidean(for rows/samples) - Clustering method:
completelinkage
If you changed these in heatmaply, mirror those parameters in your dist() and hclust() calls. Let’s start with a reproducible example using your dataset structure:
# Load your dataset (replace with your actual S) # Example: S has rows as samples, columns as features, with t1 included set.seed(123) S <- data.frame( Feature1 = rnorm(10), Feature2 = rnorm(10), Feature3 = rnorm(10) ) rownames(S) <- paste0("Sample", 1:10) S <- rbind(S, t1 = c(0.5, 0.6, 0.4)) # Add test sample t1 # Replicate heatmaply's clustering setup dist_matrix <- dist(S, method = "euclidean") # Row-wise distance (samples) hc <- hclust(dist_matrix, method = "complete") # Match heatmaply's linkage
Step 2: Understand hclust's merge Property
The merge matrix is the key here—it records every step of the clustering process:
- Each row represents one merge event (from first to last)
- Negative numbers refer to original sample indices (e.g.,
-3means the 3rd sample in your dataset) - Positive numbers refer to previously merged clusters (e.g.,
5means the cluster formed in the 5th merge step)
Since you saw t1 form a cluster in the first stage, we need to find the earliest merge that includes t1.
Step 3: Locate t1's First Merge
First, get t1's index in your dataset:
t1_idx <- which(rownames(S) == "t1")
Now scan the merge matrix to find the first row that includes -t1_idx (the negative index represents t1 as an individual sample):
# Find the first merge involving t1 target_merge <- NULL for (i in 1:nrow(hc$merge)) { if (-t1_idx %in% hc$merge[i, ]) { target_merge <- hc$merge[i, ] break } }
Step 4: Extract All Samples in the Initial Cluster
If your initial cluster is just t1 + one other sample, target_merge will have two negative numbers (e.g., -11 and -7). If it’s t1 + a small pre-formed cluster (positive number), we need a quick recursive function to unpack all samples in that sub-cluster:
# Recursive function to unpack all samples from a merge entry get_all_cluster_samples <- function(merge_entry, full_merge_matrix) { samples <- c() for (entry in merge_entry) { if (entry < 0) { samples <- c(samples, abs(entry)) # Add individual sample index } else { # Recursively unpack the pre-formed cluster samples <- c(samples, get_all_cluster_samples(full_merge_matrix[entry, ], full_merge_matrix)) } } return(unique(samples)) } # Get all indices of samples in t1's initial cluster cluster_sample_indices <- get_all_cluster_samples(target_merge, hc$merge)
Step 5: Grab the Sample Names & Features
Now it’s straightforward to extract the data you need:
# Get similar sample names similar_sample_names <- rownames(S)[cluster_sample_indices] # Extract their feature values similar_sample_features <- S[cluster_sample_indices, ] # Output the results cat("Samples similar to t1 (initial cluster):", paste(similar_sample_names, collapse = ", "), "\n") print(similar_sample_features)
Quick Note on order Property
You mentioned confusion about hc$order—this is just the order of samples as they appear in the heatmap’s dendrogram. It’s used for plotting, but we don’t need it here to extract the initial cluster.
内容的提问来源于stack exchange,提问作者Aslan

