You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:在R语言中移除数据框内的特定重复记录

解决R数据框的条件去重需求

修正后的原始数据

原代码存在语法错误,先修正为可运行版本:

snp_id <- c("chr16-54319851-C-A","chr16-54319851-C-A","chr16-54319851-C-A","chr16-54319851-C-A","chr10-100003732-A-G","")
AF_total <- c("-","-",0.1,0.1,"-","-")
df <- data.frame(snp_id, AF_total, stringsAsFactors = FALSE)

需求回顾

  1. 若snp_id重复且对应AF_total包含非-值:删除所有AF_total为-的行,再对非-的行去重
  2. 若snp_id重复但对应AF_total全为-:仅保留一行

解决方案(dplyr版)

用dplyr的分组筛选逻辑更清晰:

library(dplyr)

df_cleaned <- df %>%
  # 过滤空的snp_id(不需要可删除此行)
  filter(snp_id != "") %>%
  group_by(snp_id) %>%
  # 标记当前组是否存在非"-"的AF值
  mutate(has_non_dash = any(AF_total != "-")) %>%
  # 按条件筛选行
  filter(if(has_non_dash) AF_total != "-" else TRUE) %>%
  # 每个snp_id仅保留一行
  distinct(snp_id, .keep_all = TRUE) %>%
  ungroup() %>%
  # 删除辅助列
  select(-has_non_dash)

print(df_cleaned)

运行结果

snp_id AF_total
1 chr16-54319851-C-A      0.1
2 chr10-100003732-A-G        -

基础R实现(无需额外包)

如果不想依赖dplyr,可以用基础R的分组处理:

# 过滤空snp_id
df <- df[df$snp_id != "", ]

# 按snp_id分组处理每个子集
df_cleaned <- do.call(rbind, lapply(split(df, df$snp_id), function(sub_df) {
  has_non_dash <- any(sub_df$AF_total != "-")
  if(has_non_dash) {
    # 保留非"-"的行并去重
    unique(sub_df[sub_df$AF_total != "-", ])
  } else {
    # 仅保留第一行
    sub_df[1, ]
  }
}))

# 重置行名
rownames(df_cleaned) <- NULL

print(df_cleaned)

内容的提问来源于stack exchange,提问作者el rom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 03:46:08