You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在ggplot中仅分散重叠点且不引入随机噪声?

解决ggplot中仅分散重叠点、保留非重叠点居中的问题

需求:x轴为分类变量,y轴为连续变量,绘制带均值线、误差棒的散点图时,C、D组的点出现重叠,要求仅对重叠点进行均匀分散,非重叠点保持居中;不使用jitter(随机抖动)、position_dodge(会应用到所有类别),也不使用geom_dotplot(会丢失y值细节),同时保留y值精度。

原始示例代码

library(tidyverse)

A <- c(5.1, 5.2, 4.8)
B <- c(1.3, 2.8, 3.2)
C <- c(4.5, 4.5, 4.5)
D <- c(8.9, 7.6, 7.6)

example <- data.frame(A, B, C, D) %>%
              pivot_longer(c(A,B,C, D),
                           names_to = "Type", 
                           values_to = "Value", 
                           cols_vary = "slowest")

ggplot(example, aes(x = Type, y = Value, fill = Type)) +
  stat_summary(fun = "mean", 
               colour = "black", 
               size = 0.3,
               width = 0.4,
               geom = "crossbar") +
  stat_summary(fun.data = mean_sdl, 
               fun.args = list(mult = 1), 
               geom = "errorbar",
               linewidth = 0.8,
               width = 0.3,
               colour = "black") +
  geom_point(size = 3,
             shape = 21,
             colour = "black",
             stroke = 1)

解决方案代码

核心思路是先对数据预处理,仅给同一分类下y值重复的点添加x轴方向的固定偏移量,非重复点保持原x位置不变:

library(tidyverse)

A <- c(5.1, 5.2, 4.8)
B <- c(1.3, 2.8, 3.2)
C <- c(4.5, 4.5, 4.5)
D <- c(8.9, 7.6, 7.6)

example <- data.frame(A, B, C, D) %>%
  pivot_longer(c(A,B,C, D),
               names_to = "Type", 
               values_to = "Value", 
               cols_vary = "slowest") %>%
  # 按分类和y值分组,计算每个组内的点数量,生成偏移量
  group_by(Type, Value) %>%
  mutate(
    n = n(),
    # 仅当组内点数>1时,生成对称的偏移量,系数0.1可根据点大小调整
    x_offset = ifelse(n > 1, 
                      seq(from = -(n-1)/2, to = (n-1)/2, length.out = n) * 0.1,
                      0)
  ) %>%
  ungroup() %>%
  # 将分类变量转为数值,方便添加偏移
  mutate(x_num = as.numeric(Type))

ggplot(example, aes(y = Value, fill = Type)) +
  # 均值线和误差棒仍基于原分类位置绘制
  stat_summary(aes(x = Type),
               fun = "mean", 
               colour = "black", 
               size = 0.3,
               width = 0.4,
               geom = "crossbar") +
  stat_summary(aes(x = Type),
               fun.data = mean_sdl, 
               fun.args = list(mult = 1), 
               geom = "errorbar",
               linewidth = 0.8,
               width = 0.3,
               colour = "black") +
  # 散点使用添加偏移后的数值x轴
  geom_point(aes(x = x_num + x_offset),
             size = 3,
             shape = 21,
             colour = "black",
             stroke = 1) +
  # 将x轴刻度改回原分类名称
  scale_x_continuous(breaks = 1:4, labels = unique(example$Type)) +
  labs(x = "Type")

说明

  • 偏移量系数0.1可根据点的大小和分类间距调整,避免点超出分类范围
  • 均值线和误差棒仍基于原分类位置绘制,确保统计信息的准确性
  • 仅对同一分类下y值重复的点应用偏移,非重复点保持居中,完全符合需求

内容的提问来源于stack exchange,提问作者Jo-Maree Courtney

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 12:45:54