如何在分组直方图中添加整体直方图并统一坐标轴?
问题描述
我有一个包含多列的数据集,其中一列是取值为1、2、3的分类变量。我想用ggplot绘制3个分组直方图垂直排列,并在第三个直方图下方添加不区分分类变量的整体直方图。目前编写的代码如下:
#The name of the dataset is dataset.clust dataset.clust$cluster = factor(dataset.clust$cluster) library(patchwork) library(ggplot2) dataset.clust$cluster = factor(dataset.clust$cluster) hist_plot1 = ggplot(dataset.clust, aes(x = population, fill = cluster)) + geom_histogram(position = "identity", alpha = 0.7) + facet_wrap(~ cluster, nrow = 3) + scale_fill_manual(values = c("red", "green", "blue")) + guides(fill = guide_legend(title = "Cluster")) + ggtitle("Grouped Histograms for population") # Second Histogram hist_plot2 = ggplot(dataset.clust, aes(x = population)) + geom_histogram(fill = "yellow", color = "black") # Combine histograms and arrange vertically combined_plot = hist_plot1 / hist_plot2 print(combined_plot)
生成的图中四个子图的x轴和y轴范围不一致,请问如何调整让它们的x轴和y轴保持一致?
解决方案
要让四个图的x轴和y轴保持一致,核心是统一轴范围、分箱宽度,并通过布局对齐确保视觉统一,具体调整如下:
1. 统一轴范围
先计算population列的极值,以及所有直方图的最大计数,然后在两个绘图对象中强制设置相同的轴范围:
# 计算统一的x轴范围 x_min <- min(dataset.clust$population, na.rm = TRUE) x_max <- max(dataset.clust$population, na.rm = TRUE) # 计算所有分组及整体直方图的最大y轴计数 y_max <- max( ggplot_build(hist_plot1)$data[[1]]$count, ggplot_build(hist_plot2)$data[[1]]$count, na.rm = TRUE )
2. 统一分箱宽度
为避免自动分箱导致柱子宽度差异,在geom_histogram中显式指定binwidth参数(数值根据数据分布调整):
# 示例分箱宽度,可根据你的数据修改 bin_width <- 5
3. 调整绘图对象并对齐
修改两个绘图对象的轴设置,再用patchwork的plot_layout实现轴对齐:
# 调整分组直方图 hist_plot1 <- hist_plot1 + xlim(x_min, x_max) + ylim(0, y_max) + geom_histogram(position = "identity", alpha = 0.7, binwidth = bin_width) + theme(axis.title.x = element_blank()) # 隐藏分组图x轴标题,避免重复 # 调整整体直方图 hist_plot2 <- hist_plot2 + xlim(x_min, x_max) + ylim(0, y_max) + geom_histogram(fill = "yellow", color = "black", binwidth = bin_width) + ggtitle("Overall Histogram for population") # 组合并强制对齐 combined_plot <- hist_plot1 / hist_plot2 + plot_layout(align = "hv") # 同时水平、垂直对齐轴
完整代码
dataset.clust$cluster = factor(dataset.clust$cluster) library(patchwork) library(ggplot2) # 计算统一参数 x_min <- min(dataset.clust$population, na.rm = TRUE) x_max <- max(dataset.clust$population, na.rm = TRUE) bin_width <- 5 # 根据数据分布调整 # 绘制分组直方图 hist_plot1 = ggplot(dataset.clust, aes(x = population, fill = cluster)) + geom_histogram(position = "identity", alpha = 0.7, binwidth = bin_width) + facet_wrap(~ cluster, nrow = 3) + scale_fill_manual(values = c("red", "green", "blue")) + guides(fill = guide_legend(title = "Cluster")) + ggtitle("Grouped Histograms for population") + xlim(x_min, x_max) + ylim(0, max(ggplot_build(.)$data[[1]]$count, na.rm = TRUE)) + theme(axis.title.x = element_blank()) # 绘制整体直方图 hist_plot2 = ggplot(dataset.clust, aes(x = population)) + geom_histogram(fill = "yellow", color = "black", binwidth = bin_width) + ggtitle("Overall Histogram for population") + xlim(x_min, x_max) + ylim(0, max(ggplot_build(hist_plot1)$data[[1]]$count, ggplot_build(.)$data[[1]]$count, na.rm = TRUE)) # 组合对齐并输出 combined_plot = hist_plot1 / hist_plot2 + plot_layout(align = "hv") print(combined_plot)
内容的提问来源于stack exchange,提问作者Billy
相关产品推荐
相关产品推荐

