R语言superheat包:聚类树枝着色/加矩形及树状图复用需求
Great questions—let's break down how to solve both of your superheat needs step by step, leveraging its integration with ggdendro and R's clustering tools:
Superheat uses ggdendro under the hood for dendrograms, so we can manually replicate its clustering logic, customize the dendrogram, and pass it back into the plot.
Step 1: Replicate superheat's Column Clustering
First, we need to mirror how superheat generates the column dendrogram when col.dendrogram = TRUE and scale = TRUE:
library(superheat) library(ggdendro) library(dendextend) library(tidyverse) # Replicate superheat's column scaling and clustering scaled_mtcars <- scale(mtcars) col_dist <- dist(t(scaled_mtcars)) # Calculate distance between columns (transpose to use row distance) col_hclust <- hclust(col_dist) # Define your cluster count k (use methods like gap statistic or domain knowledge to pick k) k <- 3 col_clusters <- cutree(col_hclust, k = k)
Step 2: Color Dendrogram Branches
Use dendextend to color branches by cluster, then convert it to a ggplot object that superheat can use:
# Convert hclust to dendrogram and color branches col_dend <- as.dendrogram(col_hclust) %>% color_branches(k = k) # Generate ggdendro-compatible data for plotting dend_data <- dendro_data(col_dend, type = "rectangle") # Build custom colored dendrogram plot colored_dend_plot <- ggplot() + geom_segment(data = segment(dend_data), aes(x = x, y = y, xend = xend, yend = yend, color = color)) + scale_color_identity() + # Keep the branch colors we defined theme_dendro() # Use ggdendro's minimal theme # Pass the custom dendrogram to superheat superheat(mtcars, scale = TRUE, left.label = "none", col.dendrogram = colored_dend_plot, # Use our colored dendrogram legend = FALSE )
Step 3: Add Cluster Rectangles (Alternative to Branch Coloring)
If you prefer highlighting clusters with rectangles instead, calculate the x-axis bounds for each cluster and add them to the dendrogram plot:
# Get leaf positions from the dendrogram data leaf_positions <- dend_data$labels$x names(leaf_positions) <- dend_data$labels$label # Calculate x ranges for each cluster cluster_bounds <- map_df(unique(col_clusters), function(clust) { cluster_leaves <- names(col_clusters[col_clusters == clust]) cluster_x <- leaf_positions[cluster_leaves] tibble( xmin = min(cluster_x) - 0.5, xmax = max(cluster_x) + 0.5, ymin = 0, ymax = max(dend_data$segment$y), # Match the dendrogram's maximum height cluster = clust ) }) # Build dendrogram with cluster rectangles rect_dend_plot <- ggplot() + geom_segment(data = segment(dend_data), aes(x = x, y = y, xend = xend, yend = yend)) + geom_rect(data = cluster_bounds, aes(xmin = xmin, xmax = xmax, ymin = ymin, ymax = ymax), fill = scales::hue_pal()(k), alpha = 0.2) + # Semi-transparent rectangles theme_dendro() # Use in superheat superheat(mtcars, scale = TRUE, left.label = "none", col.dendrogram = rect_dend_plot, legend = FALSE )
To apply the same column order and dendrogram to a new dataset with identical variables, we just need to extract the cluster order from the original hclust object and enforce it on the new data.
Step 1: Extract the Original Column Order
Pull the ordered column names from the original clustering:
# Get the column order from the original hclust object original_col_order <- col_hclust$order original_col_names <- colnames(mtcars)[original_col_order]
Step 2: Apply the Order to the New Dataset
Use the extracted order to align the new dataset's columns, then pass the original dendrogram to superheat:
# Example new dataset (same variables as mtcars) set.seed(123) new_mtcars <- mtcars %>% mutate(across(everything(), ~ .x + rnorm(nrow(mtcars), 0, 0.1))) # Option 1: Pre-scale and reorder the new data, then plot scaled_new_mtcars <- scale(new_mtcars)[, original_col_names] superheat(scaled_new_mtcars, left.label = "none", col.dendrogram = colored_dend_plot, # Reuse the colored dendrogram legend = FALSE, scale = FALSE # Already scaled manually ) # Option 2: Let superheat handle scaling, but enforce column order superheat(new_mtcars, scale = TRUE, left.label = "none", col.dendrogram = colored_dend_plot, legend = FALSE, column.order = original_col_order # Force columns to match original cluster order )
This ensures the new dataset's columns are arranged exactly like the original, and the dendrogram will align perfectly with the heatmap.
内容的提问来源于stack exchange,提问作者Gustavo

