You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R数据框中筛选重复行:保留Detected为Present的条目

在R中按指定规则筛选重复行

需求:现有一个R数据框,部分variable列存在重复值,需要保留每组重复行中Detected为Present的记录;无重复的行直接保留原数据。

输入数据

# 数据结构
df <- structure(list(variable = c("CPCT02030264T", "DRUP01090004T", 
"DRUP01160001T", "DRUP01090004T", "DRUP01160001T"), value = c(0, 
0, 0, 22.2109790032886, 23.0078421452062), Detected = c("Absent", 
"Absent", "Absent", "Present", "Present"), Method = c("DNA", 
"DNA", "DNA", "RNA", "RNA")), class = "data.frame", row.names = c(39529L, 
145909L, 146881L, 304365L, 305283L))

# 打印输入数据
print(df)

输入数据预览:

variable    value Detected Method
39529  CPCT02030264T  0.00000   Absent    DNA
145909 DRUP01090004T  0.00000   Absent    DNA
146881 DRUP01160001T  0.00000   Absent    DNA
304365 DRUP01090004T 22.21098  Present    RNA
305283 DRUP01160001T 23.00784  Present    RNA

解决方案

方法1:使用dplyr包(推荐,代码简洁易读)

先安装并加载dplyr,然后按variable分组,对每组按Detected的优先级排序(把Present排在最前面),最后取每组的第一行:

library(dplyr)

result <- df %>%
  group_by(variable) %>%
  arrange(desc(Detected == "Present")) %>%  # 将Present的行排到每组首位
  slice_head(n = 1) %>%  # 提取每组第一行
  ungroup()

print(result)

方法2:Base R实现(无需额外包)

通过给Detected列赋予权重,按variable分组筛选权重最大的行:

# 给Detected赋值权重:Present=1,Absent=0
df$priority <- ifelse(df$Detected == "Present", 1, 0)

# 按variable分组,取每组priority最大的行(若有多个,取第一个出现的)
result <- df[ave(df$priority, df$variable, FUN = function(x) x == max(x)) == 1, ]
# 移除临时添加的priority列
result <- result[, !names(result) %in% "priority"]

print(result)

输出结果

两种方法都会得到如下目标结果:

variable    value Detected Method
39529  CPCT02030264T  0.00000   Absent    DNA
304365 DRUP01090004T 22.21098  Present    RNA
305283 DRUP01160001T 23.00784  Present    RNA

内容的提问来源于stack exchange,提问作者user2300940

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 02:10:38