如何在R语言中清理DataFrame的NA并合并受访者评分数据?
R语言数据框行合并与缺失值处理
现有记录受访者打分的数据框,每个受访者对应两行记录,score_1和score_2分别在不同行存有效值,其余为NA。需要将每个受访者的记录合并为一行,移除所有NA值,删除profile列,最终每个id对应一行完整数据。
原始数据
df <- data.frame( score_1 = c(1, NA, 4, NA, 9, NA, 12, NA), score_2 = c(NA, 4, NA, 29, NA, 12, NA, 9), profile = c(1, 2, 1, 2, 1, 2, 1, 2), id = c(1, 1, 2, 2, 3, 3, 4, 4), id_attribute = c(2, 2, 4, 4, 8, 8, 5, 5) )
期望结果
df_final <- data.frame( score_1 = c(1, 4, 9, 12), score_2 = c(4, 29, 12, 9), id = c(1, 2, 3, 4), id_attribute = c(2, 4, 8, 5) )
解决方案
方法1:dplyr分组聚合
按id和id_attribute分组,提取每组内非NA的有效值:
library(dplyr) df_final <- df %>% group_by(id, id_attribute) %>% summarize( score_1 = first(na.omit(score_1)), score_2 = first(na.omit(score_2)), .groups = "drop" )
na.omit过滤NA值,first取分组内唯一非NA值(每组仅一个有效值,用last也可),.groups = "drop"取消分组状态返回普通数据框。
方法2:tidyr填充去重
先分组填充缺失值,再保留唯一行:
library(tidyr) library(dplyr) df_final <- df %>% group_by(id, id_attribute) %>% fill(c(score_1, score_2), .direction = "downup") %>% distinct(id, id_attribute, score_1, score_2, .keep_all = FALSE)
fill(..., .direction = "downup")同时上下填充NA,确保同组两行都有完整score值,distinct保留唯一行完成合并。
方法3:Base R原生实现
用aggregate函数分组提取非NA值:
df_final <- aggregate( cbind(score_1, score_2) ~ id + id_attribute, data = df, FUN = function(x) x[!is.na(x)] )
自定义函数直接提取每组score列的非NA值,完成行合并。
内容的提问来源于stack exchange,提问作者juanjedi
相关产品推荐
相关产品推荐

