You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用UpSetR绘制跨物种共有基因数量的可视化图?

解决跨物种共有基因数量的UpSet图绘制问题

嘿,这个需求其实核心就是把数据结构完全反转——从「基因对应物种」变成「物种对应基因」,然后就能用你熟悉的UpSetR来画图啦!我给你一步步拆解操作:

第一步:整理原始数据

首先把你所有的基因向量打包成一个命名列表,这样后续处理更方便:

# 你的示例数据(实际替换成你20+个基因的向量)
gene1 <- c("Panda", "Dog", "Chicken")
gene2 <- c("Human", "Panda", "Dog")
gene3 <- c("Human", "Panda", "Chicken")

# 把所有基因向量放进命名列表,名称就是基因名
gene_species_list <- list(
  gene1 = gene1,
  gene2 = gene2,
  gene3 = gene3
  # 这里继续添加你的其他基因向量
)

第二步:反转数据结构(物种→基因)

我们需要把「每个基因包含哪些物种」转换成「每个物种包含哪些基因」,用tidyverse工具链处理最顺手:

library(tidyverse)

# 先转成长格式数据框:每行是「物种-基因」对应关系
long_df <- stack(gene_species_list) %>% 
  rename(species = values, gene = ind) %>% 
  distinct()  # 确保没有重复的物种-基因对(如果原始数据有重复的话)

# 再转换为「物种为名称,对应基因为元素」的列表
species_gene_list <- long_df %>% 
  group_by(species) %>% 
  summarize(genes = list(gene)) %>% 
  deframe()  # 把分组后的结果转成命名列表

第三步:用UpSetR绘制跨物种共有基因图

现在数据结构符合UpSetR的要求了,直接调用函数绘图就行:

library(UpSetR)

# 把列表转成UpSetR需要的二进制矩阵
species_matrix <- fromList(species_gene_list)

# 绘制UpSet图,自定义标签让图更清晰
upset(species_matrix,
      mainbar.y.label = "Number of Shared Genes",  # 主条形图Y轴标签
      sets.x.label = "Total Genes per Species",    # 左侧物种集合的X轴标签
      title = "Shared Genes Across Different Species",  # 图标题
      text.scale = c(1.2, 1.2, 1, 1, 1.5, 1)      # 调整文字大小,避免重叠
)

额外提示

如果你的物种数量特别多(100余种),可以通过sets参数指定要展示的物种(比如只展示基因数最多的前20个),避免图过于拥挤:

# 提取基因数最多的前10个物种
top_species <- names(sort(sapply(species_gene_list, length), decreasing = TRUE))[1:10]

# 只绘制这些物种的UpSet图
upset(species_matrix[, top_species],
      mainbar.y.label = "Number of Shared Genes",
      sets.x.label = "Total Genes per Species",
      title = "Top 10 Species: Shared Genes"
)

内容的提问来源于stack exchange,提问作者Guillermo Reales

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:25:22