You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中清理DataFrame的NA并合并受访者评分数据?

R语言数据框行合并与缺失值处理

现有记录受访者打分的数据框,每个受访者对应两行记录,score_1和score_2分别在不同行存有效值,其余为NA。需要将每个受访者的记录合并为一行,移除所有NA值,删除profile列,最终每个id对应一行完整数据。

原始数据

df <- data.frame(
  score_1 = c(1, NA, 4, NA, 9, NA, 12, NA),
  score_2 = c(NA, 4, NA, 29, NA, 12, NA, 9),
  profile = c(1, 2, 1, 2, 1, 2, 1, 2),
  id = c(1, 1, 2, 2, 3, 3, 4, 4), 
  id_attribute = c(2, 2, 4, 4, 8, 8, 5, 5)
)

期望结果

df_final <- data.frame(
  score_1 = c(1, 4, 9, 12),
  score_2 = c(4, 29, 12, 9),
  id = c(1, 2, 3, 4), 
  id_attribute = c(2, 4, 8, 5)
)

解决方案

方法1:dplyr分组聚合

按id和id_attribute分组,提取每组内非NA的有效值:

library(dplyr)

df_final <- df %>%
  group_by(id, id_attribute) %>%
  summarize(
    score_1 = first(na.omit(score_1)),
    score_2 = first(na.omit(score_2)),
    .groups = "drop"
  )

na.omit过滤NA值,first取分组内唯一非NA值(每组仅一个有效值,用last也可),.groups = "drop"取消分组状态返回普通数据框。

方法2:tidyr填充去重

先分组填充缺失值,再保留唯一行:

library(tidyr)
library(dplyr)

df_final <- df %>%
  group_by(id, id_attribute) %>%
  fill(c(score_1, score_2), .direction = "downup") %>%
  distinct(id, id_attribute, score_1, score_2, .keep_all = FALSE)

fill(..., .direction = "downup")同时上下填充NA,确保同组两行都有完整score值,distinct保留唯一行完成合并。

方法3:Base R原生实现

用aggregate函数分组提取非NA值:

df_final <- aggregate(
  cbind(score_1, score_2) ~ id + id_attribute,
  data = df,
  FUN = function(x) x[!is.na(x)]
)

自定义函数直接提取每组score列的非NA值,完成行合并。


内容的提问来源于stack exchange,提问作者juanjedi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 01:10:27