如何在R tidy数据框中按同批次同ASV对照丰度过滤样本行
过滤Tidy数据框中低于对照的样本记录
先构造符合需求的示例数据:
library(tidyverse) df <- tibble( Batch = c("B1", "B1", "B2", "B2"), Sample = c("Con1", "Sample21", "Con2", "Sample12"), Type = c("Con", "S", "Con", "S"), ASV = c("ASV1", "ASV1", "ASV2", "ASV2"), Counts = c(100, 80, 50, 60) )
核心解决方案
通过分组生成对照值列,再实现精准过滤:
filtered_df <- df %>% # 按批次和ASV分组,确保仅对比同批次同ASV的对照样本 group_by(Batch, ASV) %>% # 提取当前组内Type以Con开头的对照样本Counts值 mutate(con_counts = Counts[grepl("^Con", Type)]) %>% # 过滤规则:保留对照样本,或S样本中Counts不低于对照的记录 filter(!Type %in% "S" | (Type == "S" & Counts >= con_counts)) %>% ungroup() %>% # 移除临时生成的对照值列(可选操作) select(-con_counts)
运行后得到的结果:
# A tibble: 3 × 5 Batch Sample Type ASV Counts <chr> <chr> <chr> <chr> <dbl> 1 B1 Con1 Con ASV1 100 2 B2 Con2 Con ASV2 50 3 B2 Sample12 S ASV2 60
其中B1批次的Sample21因Counts低于对照被移除,B2批次的Sample12因Counts高于对照被保留,完全匹配需求。
处理多对照样本的情况
如果同批次同ASV下存在多个Con开头的对照样本,可将对照值改为取均值:
filtered_df <- df %>% group_by(Batch, ASV) %>% # 取所有对照样本的Counts均值作为参照 mutate(con_counts = mean(Counts[grepl("^Con", Type)], na.rm = TRUE)) %>% filter(!Type %in% "S" | (Type == "S" & Counts >= con_counts)) %>% ungroup() %>% select(-con_counts)
内容的提问来源于stack exchange,提问作者Darren
相关产品推荐
相关产品推荐

