使用dplyr按条件查找含基因A与不含A分组间的特有基因
解决思路
- 先识别出所有含基因A的分组、不含基因A的分组
- 提取不含A的分组中所有出现过的基因作为排除列表
- 从含A分组的基因列表中,排除掉基因A本身和排除列表中的基因,即为符合要求的结果
1. tidyverse实现(易读性高)
library(dplyr) # 提取含基因A的分组 group_with_A <- df %>% filter(gene == "A") %>% pull(group) %>% unique() # 提取不含A的分组中所有出现过的基因(黑名单) gene_blacklist <- df %>% filter(!group %in% group_with_A) %>% pull(gene) %>% unique() # 筛选符合要求的结果 result <- df %>% filter(group %in% group_with_A, # 仅保留含A的分组 gene != "A", # 排除A本身 !gene %in% gene_blacklist) # 排除在非A分组出现过的基因 result
输出结果和预期一致:
gene group 1 F group1 2 E group2
2. base R实现(无需额外安装包)
# 提取含基因A的分组 group_with_A <- unique(df[df$gene == "A", "group"]) # 提取基因黑名单 gene_blacklist <- unique(df[!df$group %in% group_with_A, "gene"]) # 筛选结果 result <- df[df$group %in% group_with_A & df$gene != "A" & !df$gene %in% gene_blacklist, ]
3. data.table实现(适合超大规模dataframe,效率最优)
library(data.table) setDT(df) # 转换为data.table格式 group_with_A <- unique(df[gene == "A", group]) gene_blacklist <- unique(df[!group %in% group_with_A, gene]) result <- df[group %in% group_with_A & gene != "A" & !gene %in% gene_blacklist]
内容的提问来源于stack exchange,提问作者LDT
相关产品推荐
相关产品推荐

