ggplot分组直方图X轴高值区域条形不可见的解决方法
优化ggplot2直方图X轴高值区域显示的方案
你的问题核心是数据集中在X轴低值区间,高值区间数据稀疏,导致默认等宽分箱下高值区条形被过度压缩。以下是几个贴合目标图效果的优化方案:
方案1:X轴对数变换
通过对数变换压缩高值区间的刻度间距,既保留整体分布,又让高值区条形清晰可见,适配偏态分布的数据:
library(ggplot2) library(dplyr) graph2 |> filter(!is.na(healthy)) |> ggplot(aes(x = total_visits + 1, fill = as.factor(healthy))) + # +1避免0值对数计算报错 geom_histogram(aes(y = after_stat(count / sum(count))), alpha = 0.6, color = "white", position = 'identity', breaks = exp(seq(log(1), log(101), length.out = 20))) + # 对数间距分箱 scale_x_log10(breaks = c(1, 5, 10, 20, 30, 50, 100), labels = c(0, 5, 10, 20, 30, 50, 100)) + # 标签映射回原始数值 scale_fill_manual(labels = c("TSCI", "SHS"), values = c("blue", "red")) + labs(fill = "", x = "total_visits")
方案2:截断X轴+嵌入高值区放大图
保留主图聚焦数据密集区,同时嵌入小图单独展示高值细节,兼顾整体与局部:
library(ggplot2) library(dplyr) library(ggpmisc) # 主图:展示0-30的核心分布 main_plot <- graph2 |> filter(!is.na(healthy)) |> ggplot(aes(x = total_visits, fill = as.factor(healthy))) + geom_histogram(aes(y = after_stat(count / sum(count))), alpha = 0.6, color = "white", position = 'identity', breaks = seq(0, 30, by = 1)) + scale_x_continuous(breaks = seq(0, 30, 5), limits = c(0, 30)) + scale_fill_manual(labels = c("TSCI", "SHS"), values = c("blue", "red")) + labs(fill = "", x = "total_visits") + theme(plot.margin = margin(5, 20, 5, 5)) # 子图:放大30-100的高值区 inset_plot <- graph2 |> filter(!is.na(healthy), total_visits > 30) |> ggplot(aes(x = total_visits, fill = as.factor(healthy))) + geom_histogram(aes(y = after_stat(count / sum(count))), alpha = 0.6, color = "white", position = 'identity', breaks = seq(30, 100, by = 5)) + scale_x_continuous(breaks = seq(30, 100, 10)) + scale_fill_manual(labels = c("TSCI", "SHS"), values = c("blue", "red")) + labs(x = "", y = "") + theme_minimal() + theme(legend.position = "none", plot.background = element_rect(fill = "white", color = "black")) # 组合主图与子图 main_plot + annotation_custom(grob = ggplotGrob(inset_plot), xmin = 20, xmax = 30, ymin = 0.08, ymax = 0.15)
方案3:自定义分箱宽度
对低值区用窄分箱保留细节,高值区用宽分箱避免条形过窄,平衡不同区间的显示效果:
library(ggplot2) library(dplyr) # 自定义分箱断点:0-20每1单位一箱,20-100每5单位一箱 custom_breaks <- c(seq(0, 20, 1), seq(25, 100, 5)) graph2 |> filter(!is.na(healthy)) |> ggplot(aes(x = total_visits, fill = as.factor(healthy))) + geom_histogram(aes(y = after_stat(count / sum(count))), alpha = 0.6, color = "white", position = 'identity', breaks = custom_breaks) + scale_x_continuous(breaks = c(seq(0, 20, 5), seq(25, 100, 10))) + scale_fill_manual(labels = c("TSCI", "SHS"), values = c("blue", "red")) + labs(fill = "")
内容的提问来源于stack exchange,提问作者user3483060
相关产品推荐
相关产品推荐

