R语言树状图:截取顶部11分支并实现叶节点均匀分布
解决树状图截取顶部分支后叶节点均匀分布的问题
问题背景
需要处理含1000余个叶节点的树状图,目标是保留顶部11个分支,同时让新生成的叶节点在树状图末端水平均匀分布。之前尝试通过裁剪x轴范围实现,但裁剪后叶节点间距仍基于原始20个节点的布局,无法达到均匀分布的效果。测试代码如下:
# Set seed for reproducibility set.seed(123) # Create a data frame with 20 rows and 25 columns of random data test_data <- as.data.frame(matrix(rnorm(20 * 25), nrow = 20, ncol = 25)) # Rename columns to stat1, stat2, ..., stat25 colnames(test_data) <- paste0("stat", 1:25) # Scale the data (mean = 0, SD = 1) scaled_data <- scale(test_data) # Calculate the distance matrix dist_matrix <- dist(scaled_data, method = "euclidean") # Perform hierarchical clustering hc <- hclust(dist_matrix, method = "ward.D2") # Convert the `hclust` object into a `phylo` object for ggtree phylo_tree <- as.phylo(hc) # Cut the hclust tree into 11 clusters clusters <- cutree(hc, k = 11) # Add cluster information to a data frame that matches the tips of the tree tip_data <- data.frame( label = phylo_tree$tip.label, # Tree tip labels cluster = clusters # Cluster assignments ) # Plot with tip coloring by cluster ggtree_plot <- ggtree(phylo_tree, layout = "rectangular", branch.length = "height") %<+% tip_data + geom_tiplab(aes(label = cluster), size = 2.5, align = TRUE, offset = 0.05) + geom_tippoint(aes(color = factor(cluster)), size = 3) + scale_color_manual(values = rainbow(11)) + # Use a palette with 11 distinct colors theme_tree2() print(ggtree_plot) ### CROP the tree to a height with 11 branches # Modify the ggtree plot to limit the x-axis range ggtree_plot <- ggtree(phylo_tree, layout = "rectangular", branch.length = "height") %<+% tip_data + scale_color_manual(values = rainbow(11)) + # Use a palette with 11 distinct colors theme_tree2() + coord_cartesian(xlim = c(0, 2.0)) # Set the x-axis range to 0-2.0 in this case with seed 123 print(ggtree_plot)
核心问题分析
直接裁剪x轴只是隐藏了部分分支,但树的底层结构仍基于原始所有叶节点,因此节点间距不会自动调整。要实现均匀分布,必须重新构建仅包含11个分支的树结构,将每个簇合并为单个叶节点。
解决方案代码
set.seed(123) library(ape) library(ggtree) library(phytools) # 生成测试数据(替换为你的1000+节点数据集即可) test_data <- as.data.frame(matrix(rnorm(20 * 25), nrow = 20, ncol = 25)) colnames(test_data) <- paste0("stat", 1:25) scaled_data <- scale(test_data) dist_matrix <- dist(scaled_data, method = "euclidean") hc <- hclust(dist_matrix, method = "ward.D2") phylo_tree <- as.phylo(hc) # 切割为11个簇 k <- 11 clusters <- cutree(hc, k = k) # 将簇信息转为命名向量,用于合并节点 cluster_vec <- clusters names(cluster_vec) <- phylo_tree$tip.label # 合并同一簇的所有叶节点,生成仅含11个节点的新树 merged_tree <- mergeTips(phylo_tree, cluster_vec, type = "clade") # 绘制新树,叶节点自动均匀分布 ggtree(merged_tree, layout = "rectangular", branch.length = "height") + geom_tiplab(aes(label = label), size = 3) + geom_tippoint(aes(color = factor(label)), size = 3) + scale_color_manual(values = rainbow(k)) + theme_tree2()
方案说明
- 合并节点:使用
phytools包的mergeTips函数,将原树中属于同一簇的所有叶节点合并为单个节点,生成仅包含11个叶节点的新树结构。 - 均匀分布:新树的布局完全基于11个节点,绘制时叶节点会自动水平均匀分布,不再受原始1000+节点的间距影响。
- 扩展性:该方法适用于任何规模的原始数据集,只需替换测试数据部分为你的实际数据即可。
内容的提问来源于stack exchange,提问作者R student
相关产品推荐
相关产品推荐

