如何按连续NA序列拆分R语言中的data.frame?
按连续NA值序列拆分DataFrame
我需要将一个DataFrame按指定列(示例中的temp1)的连续NA值序列拆分,得到仅包含这些连续NA片段的DataFrame列表,以便对每个序列单独处理。
示例原始数据
data <- data.frame(temp1=c(2,5,8,NA,NA,NA,7,4,1,3,NA,NA,1,5,NA,NA,NA,NA,9),temp2=c(1:19))
期望结果
得到由三个DataFrame组成的列表,每个DataFrame对应一段连续的NA序列:
result <- list( data.frame(temp1=c(NA,NA,NA),temp2=c(4,5,6)), data.frame(temp1=c(NA,NA),temp2=c(11,12)), data.frame(temp1=c(NA,NA,NA,NA),temp2=c(15,16,17,18)) )
解决方案
可以利用R的rle函数识别连续的NA段,再根据位置拆分DataFrame:
# 对temp1列的NA状态进行游程编码 na_rle <- rle(is.na(data$temp1)) # 筛选出对应连续NA的游程索引 na_runs <- which(na_rle$values) # 计算每个连续NA游程的起始与结束行号 run_starts <- cumsum(c(1, na_rle$lengths))[na_runs] run_ends <- run_starts + na_rle$lengths[na_runs] - 1 # 拆分DataFrame生成结果列表 result <- mapply(function(s, e) data[s:e, ], run_starts, run_ends, SIMPLIFY = FALSE)
运行后result即为所需的拆分列表,每个元素对应一段连续的NA序列DataFrame。
内容的提问来源于stack exchange,提问作者jeff6868
相关产品推荐
相关产品推荐

