R语言如何移除列表内数据框中D列含NA的同Study_ID全部行
问题说明
现有存储数据框的R列表对象my.list,列表内每个数据框包含Study_ID、B、C、D四个字段,每个Study_ID对应2条记录。需要按如下规则清洗数据:
- 若任意数据框中,同一个
Study_ID分组下存在任意一行的D列取值为NA,则将该Study_ID对应的全部记录从列表所有数据框中移除。
可复现测试数据构造代码如下:
my.list <- structure(list(S1 = structure(list(Study_ID = c(100, 100, 200, 200, 300,300,400,400), B = c(NA, 1.5, 1.8, 2.1, 3.2, 1.4, NA, 9.3), C = c("C1", "PTA", "C1", "PTA", "C1", "PTA","C1", "PTA"), D = c(0.9124, NA, 0.5571429, 0.7849462, 0.32719, NA, 0.82482, 0.284702 )), .Names = c("Study_ID", "B", "C", "D"), class = "data.frame", row.names = c("1", "2", "3", "4", "5", "6", "7", "8")), S2 = structure(list(Study_ID = c(100, 100, 200, 200, 300,300,400,400), B = c(NA, 0.7, NA, 0.45, 0.91, 0.78, 0.65, NA), C = c("C1", "PTA", "C1", "PTA", "C1", "PTA", "C1", "PTA"), D = c(0.9124, NA, 0.5571429, 0.7849462, 0.32719,0.6492, 0.82482, NA )), .Names = c("Study_ID", "B", "C", "D"), class = "data.frame", row.names = c("1", "2", "3", "4", "5", "6", "7", "8"))), .Names = c("S1", "S2"))
实现方法
核心逻辑分两步:
- 遍历列表所有数据框,汇总所有存在
D列NA值对应的Study_ID,作为需要剔除的ID集合 - 再次遍历列表所有数据框,保留
Study_ID不在剔除集合中的行即可
基础R实现(无需加载第三方包)
# 汇总需要剔除的Study_ID exclude_ids <- unique(unlist(lapply(my.list, function(df) df$Study_ID[is.na(df$D)]))) # 过滤所有数据框 result <- lapply(my.list, function(df) { res_df <- df[!df$Study_ID %in% exclude_ids, ] # 重置行名,和预期输出格式匹配 rownames(res_df) <- seq_len(nrow(res_df)) res_df })
tidyverse实现
如果习惯用tidyverse系列包操作,可以用如下代码:
library(dplyr) library(purrr) # 汇总需要剔除的Study_ID exclude_ids <- my.list %>% map_dfr(~ .x %>% filter(is.na(D)) %>% select(Study_ID)) %>% distinct() %>% pull(Study_ID) # 过滤所有数据框 result <- my.list %>% map(~ .x %>% filter(!Study_ID %in% exclude_ids) %>% `rownames<-`(seq_len(nrow(.))))
运行完成后result即为目标结果,打印输出和预期结果完全一致:
print(result)
输出:
$S1 Study_ID B C D 1 200 1.8 C1 0.5571429 2 200 2.1 PTA 0.7849462 3 400 NA C1 0.8248200 4 400 9.3 PTA 0.2847020 $S2 Study_ID B C D 1 200 NA C1 0.5571429 2 200 0.45 PTA 0.7849462 3 300 0.91 C1 0.3271900 4 300 0.78 PTA 0.6492000
内容的提问来源于stack exchange,提问作者sabc04
相关产品推荐
相关产品推荐

