You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言ggplot2绘制多X轴曼哈顿图排错与显著性位点标注实现

原代码问题排查
  • 变量未定义:代码中ylim未赋值,直接调用会导致Y轴渲染报错
  • 坐标映射错位:样例数据中的bp_cum为1开头的小数值,和测试用axis_set中亿级的染色体中心断点范围不匹配,所有位点会被挤在X轴最左端,染色体编号标签完全错位
  • 映射冗余:全局aes已将udi映射给color属性,geom_point中重复写color=as.factor(udi)会导致图例生成异常
  • 逻辑冲突:同时加载plyr和dplyr会引发分组聚合函数的优先级冲突,容易出现数据计算错误
  • 缺失标注逻辑:未筛选显著性位点、未添加标注图层,无法实现p<5e-8位点的ID和SNP名称标注
  • 轴配置不匹配:测试用axis_set仅覆盖1-4号染色体,无法满足1-22号染色体的展示需求
修正后完整可运行代码
# 加载依赖包,移除会和dplyr冲突的plyr
library(dplyr)
library(ggplot2)
library(ggrepel) # 用于标注自动避障

# 读取数据
gwas_data <- read.table("gwas_data", header = T, stringsAsFactors = F)
sig <- 5e-8

# --------------------------
# 第一步:计算正确的累计碱基位置与染色体轴配置(自动适配1-22号常染色体)
# --------------------------
# 计算每条染色体长度、X轴偏移量、标签中心位置
chr_set <- gwas_data %>%
  filter(chr %in% 1:22) %>%
  group_by(chr) %>%
  summarise(chr_len = max(bp), .groups = "drop") %>%
  arrange(chr) %>%
  mutate(
    chr_offset = cumsum(as.numeric(lag(chr_len, default = 0))),
    center = chr_offset + chr_len/2
  )

# 合并偏移量到原数据集,生成正确的bp_cum值
gwas_data <- gwas_data %>%
  filter(chr %in% 1:22) %>%
  left_join(chr_set %>% select(chr, chr_offset), by = "chr") %>%
  mutate(bp_cum = as.numeric(bp) + chr_offset)

# 生成正式使用的轴配置,替换原测试用4染色体axis_set
axis_set <- chr_set %>% select(chr, center)

# 自动计算Y轴上限,预留标注空间
ylim <- ceiling(max(-log10(gwas_data$p), na.rm = T)) + 2

# 筛选显著性位点,生成标注文本
sig_snps <- gwas_data %>%
  filter(p < sig) %>%
  mutate(label = paste0(udi, "\n", snp))

# --------------------------
# 第二步:绘制曼哈顿图
# --------------------------
manhplot <- ggplot(gwas_data, aes(x = bp_cum, y = -log10(p))) +
  # 显著性阈值虚线
  geom_hline(yintercept = -log10(sig), color = "grey40", linetype = "dashed") + 
  # 位点散点,按udi映射颜色
  geom_point(aes(color = as.factor(udi)), alpha = 0.8, size = 2) +
  # X轴配置
  scale_x_continuous(
    labels = axis_set$chr, 
    breaks = axis_set$center,
    expand = c(0.01, 0)
  ) +
  # Y轴配置
  scale_y_continuous(expand = c(0,0), limits = c(0, ylim)) +
  # UDI分组配色,可替换为scale_color_manual传入自定义色值
  scale_color_brewer(palette = "Set2", name = "UDI") +
  # 显著性位点标注,自动避障避免文字重叠
  geom_text_repel(
    data = sig_snps,
    aes(label = label),
    size = 3,
    min.segment.length = 0,
    segment.color = "grey50",
    nudge_y = 1.5
  ) +
  # 主题调整
  theme_minimal() +
  theme(
    panel.grid.minor = element_blank(),
    axis.title.x = element_text(size = 12, margin = margin(t = 10)),
    axis.title.y = element_text(size = 12, margin = margin(r = 10)),
    legend.position = "bottom"
  ) +
  labs(x = "Chromosome", y = expression(-log[10](italic(p))))

print(manhplot)
使用说明
  • 如果你已有预计算完成的1-22号染色体bp_cum和对应axis_set,可跳过第一步的位置计算步骤,直接替换对应变量即可
  • 若不需要标注自动避障效果,可将geom_text_repel替换为普通geom_text
  • 若需要实现标准曼哈顿图的染色体交替配色效果,只需将color映射改为as.factor(chr)即可
  • 可通过scale_color_manual(values = 自定义色值向量)替换UDI分组的配色方案

内容的提问来源于stack exchange,提问作者zerberus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 21:54:28