ggplot2中两组整数数据抖动点差异及分布平滑优化问题
问题解答
数据与绘图代码
数据集
mydata <- structure(list(group = c("a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "a", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b", "b"), age = c(27, 21, 27, 21, 24, 27, 27, 30, 27, 27, 27, 27, 27, 21, 27, 24, 24, 27, 21, 24, 30, 30, 24, 24, 27, 18, 27, 24, 18, 18, 9, 12, 15, 12, 15, 12, 9, 9, 12, 15, 15, 18, 18, 15, 21, 21, 15, 21, 12, 21, 15, 30, 21, 18, 21, 21, 24, 21, 24, 24, 27, 24, 18, 27, 9, 21, 27, 21, 21, 21, 27, 24, 27, 24, 30, 30, 30, 27, 27, 24, 27, 27, 24, 24, 30, 27, 27, 30, 21, 24, 21, 27, 24, 24, 24, 24, 24, 24, 24, 21, 34, 25, 27, 35, 27, 28, 32, 33, 32, 9, 9, 8, 15, 29, 30, 10, 40, 31, 27, 40, 28, 31, 17, 19, 35, 29, 23, 15, 16, 26, 27, 25, 23, 24, 25, 25, 13, 36, 25, 27, 35, 35, 24, 21, 25, 10, 23, 5, 34, 21)), row.names = c(1L, 2L, 3L, 4L, 5L, 6L, 7L, 8L, 9L, 10L, 11L, 12L, 13L, 14L, 15L, 16L, 17L, 18L, 19L, 20L, 21L, 22L, 23L, 24L, 25L, 26L, 27L, 28L, 29L, 30L, 31L, 32L, 33L, 34L, 35L, 36L, 37L, 38L, 39L, 40L, 41L, 42L, 43L, 44L, 45L, 46L, 47L, 48L, 49L, 50L, 51L, 52L, 53L, 54L, 55L, 56L, 57L, 58L, 59L, 60L, 61L, 62L, 63L, 64L, 65L, 66L, 67L, 68L, 69L, 70L, 71L, 72L, 73L, 74L, 75L, 76L, 77L, 78L, 79L, 80L, 81L, 82L, 83L, 84L, 85L, 86L, 87L, 88L, 89L, 90L, 91L, 92L, 93L, 94L, 95L, 96L, 97L, 98L, 99L, 100L, 101L, 102L, 103L, 104L, 105L, 106L, 107L, 108L, 109L, 110L, 111L, 112L, 113L, 114L, 115L, 116L, 117L, 118L, 119L, 120L, 121L, 122L, 123L, 124L, 125L, 126L, 127L, 128L, 129L, 130L, 132L, 133L, 134L, 135L, 136L, 137L, 138L, 139L, 140L, 141L, 142L, 143L, 144L, 145L, 146L, 147L, 148L, 149L, 150L, 151L), class = "data.frame")
绘图代码
library(ggplot2) library(ggdist) ggplot(mydata, aes(x = group, y = age)) + ggdist::stat_halfeye( adjust = .5, width = .6, .width = 0, justification = -.3, point_colour = NA) + geom_point( size = 1.3, alpha = .3, position = position_jitter( seed = 1, width = .1 ) )
绘图结果

现象原因
- 抖动幅度差异:group a的样本量远大于group b(98个vs53个),即使设置了相同的抖动宽度
width=.1,样本量大会让点更密集地重叠,看起来抖动幅度更小;而group b样本量小,点分布更分散,视觉上抖动幅度更大。 - 分布曲线平滑度差异:
stat_halfeye用核密度估计生成曲线,平滑度受两个因素影响:- 数据分布:group a的age高度集中在几个整数(比如24、27出现频次极高),属于离散且集中的分布,核密度估计在这些密集点附近会形成“峰值”,导致曲线不平滑;而group b的age数值更分散,覆盖的整数更多,曲线自然更平滑。
adjust参数:当前设置的adjust=.5会让核密度估计更“敏感”,放大了group a数据集中的峰值差异,加剧了曲线的不平滑。
解决方法
要让group a的分布曲线和group b一样平滑,只需增大stat_halfeye中的adjust参数值,让核密度估计更“宽泛”,抹平密集点带来的峰值:
ggplot(mydata, aes(x = group, y = age)) + ggdist::stat_halfeye( adjust = 1.5, # 增大adjust值,可根据实际效果调整为1、2等 width = .6, .width = 0, justification = -.3, point_colour = NA) + geom_point( size = 1.3, alpha = .3, position = position_jitter( seed = 1, width = .1 ) )
你可以根据视觉效果调整adjust的数值,直到group a的曲线平滑度符合预期。如果想让抖动视觉上更一致,也可以给group a设置稍大一点的抖动宽度,但核心解决方法还是调整核密度的adjust参数。
内容的提问来源于stack exchange,提问作者Patrick
相关产品推荐
相关产品推荐

