在R中按分组为数据集补充缺失horizon的空白行
按分组补全缺失Horizon的R语言实现
原始数据集
dat <- data.frame( group = c(1,1,1,1,1,1,2,2,2,2,2), horizon = c(1,3,5,6,7,10,1,3,5,9,10), value = c(1.0,0.9,0.8,0.6,0.3,0.0,0.5,0.6,0.8,0.9,0.8), other = c("a","a","a","a","a","a","b","b","b","b","b") )
需求说明
需按group分组,为每个组补全1~10的所有horizon值:
- 缺失的
horizon对应的value设为NA - 保留
other变量(每个分组的other值固定)
解决方案
方法1:tidyverse 实现(简洁高效,适合大数据集)
利用tidyr::complete()函数直接补全分组内的缺失组合:
library(tidyverse) datx <- dat %>% group_by(group, other) %>% complete(horizon = seq(1, 10, 1)) %>% ungroup() %>% arrange(group, horizon)
注:
complete()会自动为每个分组生成指定范围内的所有horizon值,缺失的value自动填充NA,同时保留分组对应的other值。
方法2:Base R 实现(无需额外安装包)
通过生成完整的分组-horizon序列,再与原始数据合并:
# 获取每个group对应的唯一other值 group_other_map <- unique(dat[, c("group", "other")]) # 生成所有group的完整horizon序列 full_seq <- expand.grid( group = group_other_map$group, horizon = seq(1, 10, 1), stringsAsFactors = FALSE ) # 关联other值并合并原始数据 full_seq <- merge(full_seq, group_other_map, by = "group", all.x = TRUE) datx <- merge(full_seq, dat, by = c("group", "horizon", "other"), all.x = TRUE) # 排序并重置行名 datx <- datx[order(datx$group, datx$horizon), ] rownames(datx) <- NULL
内容的提问来源于stack exchange,提问作者Jan
相关产品推荐
相关产品推荐

