如何更高效编写str_detect代码?精简数据提取适配UpSet图
问题需求
需要从数据框的某列提取指定字符对应的行名,每个匹配字符对应的行名存入单独的对象,用于绘制UpSet图展示各集合的重叠情况。现有代码冗长重复,需精简优化。
模拟数据(修正后可运行)
# 修正rep用法,生成100行有效数据 name <- rep(c("rs123", "rs124", "rs125", "rs126", "rs127", "rs128", "rs129", "rs130"), length.out = 100) source <- rep("dbSNP", 100) chromosome <- rep("17", 100) evidence <- rep(c("TOPMed", "Frequency", "Cited", "Frequency, Cited", "Disease", "ESP", "TOPMed", "GWAS", "1000Genome"), length.out = 100) Biomart <- data.frame(name, source, chromosome, evidence)
现有冗余代码
'1000Genome' <- Biomart[str_detect(Biomart$evidence,"1000Genomes"),]$name 'freq' <- Biomart[str_detect(Biomart$evidence,"Frequency"),]$name 'TOPMed' <- Biomart[str_detect(Biomart$evidence,"TOPMed"),]$name 'gnomAD' <- Biomart[str_detect(Biomart$evidence, "gnomAD"),]$name 'Cited' <- Biomart[str_detect(Biomart$evidence, "Cited"),]$name 'ESP' <- Biomart[str_detect(Biomart$evidence, "ESP"),]$name 'ExAC' <- Biomart[str_detect(Biomart$evidence, "ExAC"),]$name 'Phenotype_or_Disease' <- Biomart[str_detect(Biomart$evidence, "Disease"),]$name
精简方案
方案1:用循环生成集合列表
把需要匹配的字符存入向量,通过循环批量生成每个集合的行名,结果存入列表(方便统一管理和调用):
library(stringr) # 定义需要匹配的证据类型:键为集合名,值为匹配字符串 evidence_types <- list( "1000Genome" = "1000Genome", "freq" = "Frequency", "TOPMed" = "TOPMed", "gnomAD" = "gnomAD", "Cited" = "Cited", "ESP" = "ESP", "ExAC" = "ExAC", "Phenotype_or_Disease" = "Disease" ) # 循环生成每个集合的行名,存入列表 evidence_sets <- lapply(evidence_types, function(pattern) { Biomart[str_detect(Biomart$evidence, pattern), "name"] }) # 调用单个集合示例:evidence_sets$TOPMed
方案2:生成UpSet图直接可用的二进制矩阵
UpSet图通常需要每行对应一个样本,每列对应一个集合,值为1/0表示样本是否属于该集合,这种格式可直接传入UpSetR或ComplexUpset包,省去后续整理步骤:
library(stringr) library(dplyr) # 定义集合名称和对应匹配字符串 evidence_cols <- c( "1000Genome", "freq", "TOPMed", "gnomAD", "Cited", "ESP", "ExAC", "Phenotype_or_Disease" ) patterns <- c( "1000Genome", "Frequency", "TOPMed", "gnomAD", "Cited", "ESP", "ExAC", "Disease" ) # 生成二进制矩阵 upset_data <- Biomart %>% select(name) %>% mutate(across(all_of(evidence_cols), ~str_detect(Biomart$evidence, patterns[which(evidence_cols == cur_column())]))) %>% mutate(across(-name, as.integer)) # 查看结果示例 head(upset_data)
说明
- 方案1保留了「每个集合单独存储」的需求,用列表替代多个独立变量,新增集合只需修改
evidence_types,维护性更强。 - 方案2直接生成UpSet图所需格式,一步到位,是更高效的选择。
内容的提问来源于stack exchange,提问作者Ted
相关产品推荐
相关产品推荐

