在R语言中选择树状图中低于指定高度且规模大于1的聚类
Hey there! I totally get your frustration with cutree dragging in all those single-member clusters—let's fix this with a fully automated workflow, no tedious manual selection needed. Here's how to tackle it step by step:
Step 1: Run Hierarchical Clustering & Initial Cut
First, let's start with the core clustering setup (skip this if you already have your hierarchical clustering object ready):
# Replace `your_data` with your actual dataset hc <- hclust(dist(your_data))
Now use cutree with your target height threshold to get initial cluster assignments:
# Set your desired height threshold here threshold_height <- 5 clusters <- cutree(hc, h = threshold_height)
Step 2: Filter Out Single-Member Clusters
The trick here is to isolate clusters with more than one member, so we can ignore those tiny, unhelpful singleton groups:
# Calculate how many samples are in each cluster cluster_sizes <- table(clusters) # Keep only clusters that have more than 1 sample valid_cluster_ids <- as.numeric(names(cluster_sizes)[cluster_sizes > 1])
From here, you can grab the indices of all samples belonging to these valid clusters:
target_sample_indices <- which(clusters %in% valid_cluster_ids)
Step 3: Highlight the Target Clusters in the Dendrogram
You’ve got two straightforward options to make these clusters stand out—pick whichever fits your visualization needs:
Option 1: Draw Bold Boxes Around Valid Clusters
Use rect.hclust to add eye-catching borders around your target groups:
plot(hc, main = "Dendrogram with Highlighted Non-Singleton Clusters") # Draw thick red boxes around the clusters we care about rect.hclust(hc, which = valid_cluster_ids, border = "red", lwd = 2)
Option 2: Color the Sample Labels of Valid Clusters
If you prefer colored labels over boxes, use this approach to make target samples pop:
plot(hc, main = "Dendrogram with Colored Target Cluster Labels") # Set default label color to black, then update target samples to red label_colors <- rep("black", length(clusters)) label_colors[target_sample_indices] <- "red" # Add the colored labels to the dendrogram text(hc$order, labels = rownames(your_data)[hc$order], col = label_colors, cex = 0.8)
Quick Note on Why This Fixes Your Issue
cutree includes single-member clusters because it splits the dendrogram exactly at your threshold height—any branch that hasn’t merged with others by that point gets its own cluster. By filtering out clusters with size 1, we only keep the meaningful groups you want, all without manual tinkering.
内容的提问来源于stack exchange,提问作者Kate Bamford

