求助:在R语言中移除数据框内的特定重复记录
解决R数据框的条件去重需求
修正后的原始数据
原代码存在语法错误,先修正为可运行版本:
snp_id <- c("chr16-54319851-C-A","chr16-54319851-C-A","chr16-54319851-C-A","chr16-54319851-C-A","chr10-100003732-A-G","") AF_total <- c("-","-",0.1,0.1,"-","-") df <- data.frame(snp_id, AF_total, stringsAsFactors = FALSE)
需求回顾
- 若
snp_id重复且对应AF_total包含非-值:删除所有AF_total为-的行,再对非-的行去重 - 若
snp_id重复但对应AF_total全为-:仅保留一行
解决方案(dplyr版)
用dplyr的分组筛选逻辑更清晰:
library(dplyr) df_cleaned <- df %>% # 过滤空的snp_id(不需要可删除此行) filter(snp_id != "") %>% group_by(snp_id) %>% # 标记当前组是否存在非"-"的AF值 mutate(has_non_dash = any(AF_total != "-")) %>% # 按条件筛选行 filter(if(has_non_dash) AF_total != "-" else TRUE) %>% # 每个snp_id仅保留一行 distinct(snp_id, .keep_all = TRUE) %>% ungroup() %>% # 删除辅助列 select(-has_non_dash) print(df_cleaned)
运行结果
snp_id AF_total 1 chr16-54319851-C-A 0.1 2 chr10-100003732-A-G -
基础R实现(无需额外包)
如果不想依赖dplyr,可以用基础R的分组处理:
# 过滤空snp_id df <- df[df$snp_id != "", ] # 按snp_id分组处理每个子集 df_cleaned <- do.call(rbind, lapply(split(df, df$snp_id), function(sub_df) { has_non_dash <- any(sub_df$AF_total != "-") if(has_non_dash) { # 保留非"-"的行并去重 unique(sub_df[sub_df$AF_total != "-", ]) } else { # 仅保留第一行 sub_df[1, ] } })) # 重置行名 rownames(df_cleaned) <- NULL print(df_cleaned)
内容的提问来源于stack exchange,提问作者el rom
相关产品推荐
相关产品推荐

