You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更高效编写str_detect代码?精简数据提取适配UpSet图

问题需求

需要从数据框的某列提取指定字符对应的行名,每个匹配字符对应的行名存入单独的对象,用于绘制UpSet图展示各集合的重叠情况。现有代码冗长重复,需精简优化。

模拟数据(修正后可运行)

# 修正rep用法,生成100行有效数据
name <- rep(c("rs123", "rs124", "rs125", "rs126", "rs127", "rs128", "rs129", "rs130"), length.out = 100)
source <- rep("dbSNP", 100)
chromosome <- rep("17", 100)
evidence <- rep(c("TOPMed", "Frequency", "Cited", "Frequency, Cited", "Disease", "ESP", "TOPMed", "GWAS", "1000Genome"), length.out = 100)

Biomart <- data.frame(name, source, chromosome, evidence)

现有冗余代码

'1000Genome' <- Biomart[str_detect(Biomart$evidence,"1000Genomes"),]$name 
'freq' <- Biomart[str_detect(Biomart$evidence,"Frequency"),]$name
'TOPMed' <-  Biomart[str_detect(Biomart$evidence,"TOPMed"),]$name
'gnomAD' <- Biomart[str_detect(Biomart$evidence, "gnomAD"),]$name
'Cited' <- Biomart[str_detect(Biomart$evidence, "Cited"),]$name
'ESP' <- Biomart[str_detect(Biomart$evidence, "ESP"),]$name
'ExAC' <- Biomart[str_detect(Biomart$evidence, "ExAC"),]$name
'Phenotype_or_Disease' <- Biomart[str_detect(Biomart$evidence, "Disease"),]$name

精简方案

方案1:用循环生成集合列表

把需要匹配的字符存入向量,通过循环批量生成每个集合的行名,结果存入列表(方便统一管理和调用):

library(stringr)

# 定义需要匹配的证据类型:键为集合名,值为匹配字符串
evidence_types <- list(
  "1000Genome" = "1000Genome",
  "freq" = "Frequency",
  "TOPMed" = "TOPMed",
  "gnomAD" = "gnomAD",
  "Cited" = "Cited",
  "ESP" = "ESP",
  "ExAC" = "ExAC",
  "Phenotype_or_Disease" = "Disease"
)

# 循环生成每个集合的行名,存入列表
evidence_sets <- lapply(evidence_types, function(pattern) {
  Biomart[str_detect(Biomart$evidence, pattern), "name"]
})

# 调用单个集合示例:evidence_sets$TOPMed

方案2:生成UpSet图直接可用的二进制矩阵

UpSet图通常需要每行对应一个样本,每列对应一个集合,值为1/0表示样本是否属于该集合,这种格式可直接传入UpSetR或ComplexUpset包,省去后续整理步骤:

library(stringr)
library(dplyr)

# 定义集合名称和对应匹配字符串
evidence_cols <- c(
  "1000Genome", "freq", "TOPMed", "gnomAD", 
  "Cited", "ESP", "ExAC", "Phenotype_or_Disease"
)
patterns <- c(
  "1000Genome", "Frequency", "TOPMed", "gnomAD", 
  "Cited", "ESP", "ExAC", "Disease"
)

# 生成二进制矩阵
upset_data <- Biomart %>%
  select(name) %>%
  mutate(across(all_of(evidence_cols), ~str_detect(Biomart$evidence, patterns[which(evidence_cols == cur_column())]))) %>%
  mutate(across(-name, as.integer))

# 查看结果示例
head(upset_data)

说明

  • 方案1保留了「每个集合单独存储」的需求,用列表替代多个独立变量,新增集合只需修改evidence_types,维护性更强。
  • 方案2直接生成UpSet图所需格式,一步到位,是更高效的选择。

内容的提问来源于stack exchange,提问作者Ted

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 11:35:17