You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用dplyr筛选满足组内样本丰度分布条件的列?

基于组内样本丰度条件筛选DataFrame列(dplyr实现)

输入数据

先构造可复现的输入数据:

library(tidyverse)

df <- tibble(
  sample = paste0("Sample", 1:12),
  ColumnA = c(0.05, 0.2, 0.04, 1e-04, 0.03, 0.05, 0.006, 8e-04, 4e-04, 0.008, 0.02, 0.035),
  ColumnB = c(0.00042, 0.00084, 0.00045, 0.13, 0.0058, 0.0066, 0.022, 0.101, 2e-04, 5e-04, 8.6e-05, 4.3e-05),
  ColumnC = c(0.12, 0.223, 0.063, 0.013, 3e-04, 0.00021, 3.2e-06, 0.00063, 0.000233, 0.00075, 0.00076, 0.00044),
  group = rep(c("X", "Y", "Z"), each = 4)
) %>%
  column_to_rownames("sample")

筛选规则

保留满足以下条件的列:至少存在一个组,该组内有≥3个样本的丰度值>0.01。示例中符合条件的为ColumnA和ColumnC。

dplyr解决方案

# 第一步:计算每个列是否符合筛选条件
col_keep <- df %>%
  # 转换为长格式,便于按列和组统计
  pivot_longer(cols = starts_with("Column"), names_to = "column", values_to = "value") %>%
  # 按列和组分组,统计每组中值>0.01的样本数
  group_by(column, group) %>%
  summarise(gt_count = sum(value > 0.01), .groups = "drop_last") %>%
  # 对每个列判断是否存在至少一个组满足"≥3个样本>0.01"
  summarise(keep = any(gt_count >= 3)) %>%
  # 提取为命名逻辑向量,列名为键,是否保留为值
  deframe()

# 第二步:筛选符合条件的列
filtered_df <- df %>%
  select(names(col_keep)[col_keep], group)

结果验证

运行上述代码后,filtered_df的结构如下:

> filtered_df
         ColumnA ColumnC group
Sample1     0.05    0.12     X
Sample2      0.2   0.223     X
Sample3     0.04   0.063     X
Sample4    1e-04   0.013     X
Sample5     0.03  3e-04     Y
Sample6     0.05 0.00021     Y
Sample7    0.006 3.2e-06     Y
Sample8    8e-04 0.00063     Y
Sample9    4e-04 0.000233     Z
Sample10   0.008  0.00075     Z
Sample11    0.02  0.00076     Z
Sample12   0.035  0.00044     Z

内容的提问来源于stack exchange,提问作者eraysahin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 14:02:51