You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中构建跨双轴的镜像棒棒糖图:解决数据集规模差异导致的轴扭曲问题

解决两组规模差异悬殊的棒棒糖图对比问题

看起来你遇到的核心问题是两组数据集的样本量(行数)差异太大,直接用绝对计数作图会导致其中一组的变异信号被压缩得几乎没法直观对比——这是这类对比图里很常见的坑,我给你几个可行的解决方案,优先推荐相对频率转换的思路,因为它更符合统计对比的合理性:

核心思路:把绝对计数转换成相对频率

既然两组的总样本数差了十几倍,我们不能直接比“变异的绝对数量”,应该比“单位样本的变异频率”(或者每组内变异的占比),这样两组的数值范围会拉到同一量级,坐标轴就不会扭曲了。同时保留你原来的上下布局:把case组放在Y轴正方向,control组放在负方向,用颜色区分不同变异类型。

修改后的完整代码(推荐方案)

假设你的小数据集叫plot_case,大数据集叫plot_control,代码如下:

library(tidyverse)

# 第一步:预处理两组数据,计算相对频率
# Case组:计算每个外显子-变异类型的绝对计数,再转换成「每样本的变异频率」
case_processed <- plot_case %>%
  mutate(group = "Case") %>%
  count(Exon = as.numeric(Exon), Variant_Classification, group) %>%
  # 用变异数除以组内总样本数(这里用数据集行数代表样本量)
  mutate(rel_freq = n / nrow(plot_case))

# Control组:同样计算相对频率,但取负数让它落在Y轴下方
control_processed <- plot_control %>%
  mutate(group = "Control") %>%
  count(Exon = as.numeric(Exon), Variant_Classification, group) %>%
  mutate(rel_freq = -n / nrow(plot_control))

# 合并两组数据
combined_df <- bind_rows(case_processed, control_processed)

# 第二步:绘制优化后的棒棒糖图
ggplot(combined_df, aes(x = Exon, color = Variant_Classification)) +
  # 画Y=0的分隔线,区分上下两组
  geom_hline(yintercept = 0, color = "black", linewidth = 0.8) +
  # 棒棒糖的竖线:从Y=0延伸到对应频率值
  geom_linerange(aes(ymin = 0, ymax = rel_freq), 
                 position = position_dodge(width = 0.5), linewidth = 0.8) +
  # 棒棒糖的圆点:标记对应频率的位置
  geom_point(aes(y = rel_freq), size = 3, 
             position = position_dodge(width = 0.5)) +
  # X轴设置:显示全部30个外显子的刻度
  scale_x_continuous(breaks = 1:30, expand = c(0.02, 0)) +
  # Y轴设置:显示正负频率,标签用绝对值(避免负数干扰阅读)
  scale_y_continuous(
    breaks = seq(-0.2, 0.2, by = 0.05), 
    labels = abs,
    name = "Relative Mutation Frequency (per sample)"
  ) +
  # 添加两组的标签,放在左上角位置
  annotate("text", x = 1, y = c(-0.18, 0.18), 
           label = c("Control (456 samples)", "Case (27 samples)"),
           fontface = "bold") +
  # 主题优化:让图表更简洁美观
  theme_minimal() +
  theme(
    legend.position = "bottom",
    axis.text.x = element_text(size = 10),
    panel.grid.minor = element_blank(),
    legend.title = element_text(face = "bold"),
    plot.title = element_text(hjust = 0.5, face = "bold", size = 14)
  ) +
  labs(title = "Mutation Frequency Comparison by Exon",
       color = "Variant Classification")

关键调整说明

  1. 相对频率转换:把绝对计数n转换成n / 组内总样本数,这样两组的数值范围会处于同一量级(比如case组可能是0-0.2,control组是-0.2到0),完美解决坐标轴扭曲的问题。如果你更关注变异在组内的占比,也可以把分母换成sum(n)(每组内所有变异数的总和)。
  2. Control组取负:让Control组的数值落在Y轴负方向,和Case组的正方向形成清晰的上下对比,一眼就能区分两组。
  3. 位置避让:用position_dodge(width = 0.5)确保同一外显子的不同变异类型的点不会重叠,保证可读性。
  4. 坐标轴优化:调整Y轴的刻度范围和步长,确保两组的变异点都能清晰展示;X轴保留所有30个外显子的刻度,符合你的需求。

备选方案:双Y轴(谨慎使用)

如果你坚持要用绝对计数展示,也可以用双Y轴,但双Y轴容易误导读者,一定要明确标注两组的刻度:

library(tidyverse)

# 预处理两组的绝对计数
case_count <- plot_case %>% 
  count(Exon = as.numeric(Exon), Variant_Classification) %>% 
  mutate(group = "Case")

control_count <- plot_control %>% 
  count(Exon = as.numeric(Exon), Variant_Classification) %>% 
  mutate(group = "Control")

# 双Y轴版本的棒棒糖图
ggplot() +
  # Case组:绑定右侧Y轴
  geom_linerange(data = case_count, 
                 aes(x = Exon, ymin = 0, ymax = n, color = Variant_Classification),
                 position = position_dodge(width = 0.5), linewidth = 0.8) +
  geom_point(data = case_count, 
             aes(x = Exon, y = n, color = Variant_Classification),
             position = position_dodge(width = 0.5), size = 3) +
  # Control组:绑定左侧Y轴(取负让它在下方)
  geom_linerange(data = control_count, 
                 aes(x = Exon, ymin = 0, ymax = -n, color = Variant_Classification),
                 position = position_dodge(width = 0.5), linewidth = 0.8) +
  geom_point(data = control_count, 
             aes(x = Exon, y = -n, color = Variant_Classification),
             position = position_dodge(width = 0.5), size = 3) +
  # X轴设置
  scale_x_continuous(breaks = 1:30) +
  # 双Y轴配置:左侧是Control的计数,右侧是Case的计数
  scale_y_continuous(
    name = "Control Mutation Count",
    breaks = seq(-10, 0, by = 2), labels = abs(seq(-10, 0, by = 2)),
    sec.axis = sec_axis(~., name = "Case Mutation Count", breaks = seq(0, 4, by = 1))
  ) +
  # 中间分隔线
  geom_hline(yintercept = 0, color = "black", linewidth = 0.8) +
  # 组标签
  annotate("text", x = 1, y = c(-8, 3), 
           label = c("Control (456 samples)", "Case (27 samples)"),
           fontface = "bold") +
  # 主题设置
  theme_minimal() +
  theme(legend.position = "bottom")

不过还是更推荐第一个相对频率的方案,因为它更符合统计对比的逻辑,不会让读者产生“Control组变异远多于Case组”的误解——实际上是Control组样本量更大,相对频率可能反而更低。


内容的提问来源于stack exchange,提问作者tacrolimus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:47:41