You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R tidy数据框中按同批次同ASV对照丰度过滤样本行

过滤Tidy数据框中低于对照的样本记录

先构造符合需求的示例数据:

library(tidyverse)

df <- tibble(
  Batch = c("B1", "B1", "B2", "B2"),
  Sample = c("Con1", "Sample21", "Con2", "Sample12"),
  Type = c("Con", "S", "Con", "S"),
  ASV = c("ASV1", "ASV1", "ASV2", "ASV2"),
  Counts = c(100, 80, 50, 60)
)

核心解决方案

通过分组生成对照值列,再实现精准过滤:

filtered_df <- df %>%
  # 按批次和ASV分组,确保仅对比同批次同ASV的对照样本
  group_by(Batch, ASV) %>%
  # 提取当前组内Type以Con开头的对照样本Counts值
  mutate(con_counts = Counts[grepl("^Con", Type)]) %>%
  # 过滤规则:保留对照样本,或S样本中Counts不低于对照的记录
  filter(!Type %in% "S" | (Type == "S" & Counts >= con_counts)) %>%
  ungroup() %>%
  # 移除临时生成的对照值列(可选操作)
  select(-con_counts)

运行后得到的结果:

# A tibble: 3 × 5
  Batch Sample  Type  ASV   Counts
  <chr> <chr>   <chr> <chr>  <dbl>
1 B1    Con1    Con   ASV1     100
2 B2    Con2    Con   ASV2      50
3 B2    Sample12 S     ASV2      60

其中B1批次的Sample21因Counts低于对照被移除,B2批次的Sample12因Counts高于对照被保留,完全匹配需求。

处理多对照样本的情况

如果同批次同ASV下存在多个Con开头的对照样本,可将对照值改为取均值:

filtered_df <- df %>%
  group_by(Batch, ASV) %>%
  # 取所有对照样本的Counts均值作为参照
  mutate(con_counts = mean(Counts[grepl("^Con", Type)], na.rm = TRUE)) %>%
  filter(!Type %in% "S" | (Type == "S" & Counts >= con_counts)) %>%
  ungroup() %>%
  select(-con_counts)

内容的提问来源于stack exchange,提问作者Darren

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 06:01:05