如何在ggplot中仅分散重叠点且不引入随机噪声?
解决ggplot中仅分散重叠点、保留非重叠点居中的问题
需求:x轴为分类变量,y轴为连续变量,绘制带均值线、误差棒的散点图时,C、D组的点出现重叠,要求仅对重叠点进行均匀分散,非重叠点保持居中;不使用jitter(随机抖动)、position_dodge(会应用到所有类别),也不使用geom_dotplot(会丢失y值细节),同时保留y值精度。
原始示例代码
library(tidyverse) A <- c(5.1, 5.2, 4.8) B <- c(1.3, 2.8, 3.2) C <- c(4.5, 4.5, 4.5) D <- c(8.9, 7.6, 7.6) example <- data.frame(A, B, C, D) %>% pivot_longer(c(A,B,C, D), names_to = "Type", values_to = "Value", cols_vary = "slowest") ggplot(example, aes(x = Type, y = Value, fill = Type)) + stat_summary(fun = "mean", colour = "black", size = 0.3, width = 0.4, geom = "crossbar") + stat_summary(fun.data = mean_sdl, fun.args = list(mult = 1), geom = "errorbar", linewidth = 0.8, width = 0.3, colour = "black") + geom_point(size = 3, shape = 21, colour = "black", stroke = 1)
解决方案代码
核心思路是先对数据预处理,仅给同一分类下y值重复的点添加x轴方向的固定偏移量,非重复点保持原x位置不变:
library(tidyverse) A <- c(5.1, 5.2, 4.8) B <- c(1.3, 2.8, 3.2) C <- c(4.5, 4.5, 4.5) D <- c(8.9, 7.6, 7.6) example <- data.frame(A, B, C, D) %>% pivot_longer(c(A,B,C, D), names_to = "Type", values_to = "Value", cols_vary = "slowest") %>% # 按分类和y值分组,计算每个组内的点数量,生成偏移量 group_by(Type, Value) %>% mutate( n = n(), # 仅当组内点数>1时,生成对称的偏移量,系数0.1可根据点大小调整 x_offset = ifelse(n > 1, seq(from = -(n-1)/2, to = (n-1)/2, length.out = n) * 0.1, 0) ) %>% ungroup() %>% # 将分类变量转为数值,方便添加偏移 mutate(x_num = as.numeric(Type)) ggplot(example, aes(y = Value, fill = Type)) + # 均值线和误差棒仍基于原分类位置绘制 stat_summary(aes(x = Type), fun = "mean", colour = "black", size = 0.3, width = 0.4, geom = "crossbar") + stat_summary(aes(x = Type), fun.data = mean_sdl, fun.args = list(mult = 1), geom = "errorbar", linewidth = 0.8, width = 0.3, colour = "black") + # 散点使用添加偏移后的数值x轴 geom_point(aes(x = x_num + x_offset), size = 3, shape = 21, colour = "black", stroke = 1) + # 将x轴刻度改回原分类名称 scale_x_continuous(breaks = 1:4, labels = unique(example$Type)) + labs(x = "Type")
说明
- 偏移量系数
0.1可根据点的大小和分类间距调整,避免点超出分类范围 - 均值线和误差棒仍基于原分类位置绘制,确保统计信息的准确性
- 仅对同一分类下y值重复的点应用偏移,非重复点保持居中,完全符合需求
内容的提问来源于stack exchange,提问作者Jo-Maree Courtney
相关产品推荐
相关产品推荐

