使用lapply+split处理data.table时rbindlist报错求助
data.table结合lapply/split后绑定异常问题
问题现象
使用data.table_1.14.2版本时,通过split拆分data.table后用lapply处理,再用rbindlist或do.call(rbind)合并,无法得到包含原数据的完整结果:
rbindlist(list(1:3, 4:6))直接报错,提示输入不是数据框/表/列表- 对示例数据
temp_data处理时:test1执行报错,无法生成对象test2仅返回新计算的cases_daily列构成的矩阵,原数据全部丢失test3的lapply结果仅为cases_daily的列向量
- 但手动拆分每个分组处理后,用
rbindlist绑定能得到正常结果
代码示例
# 报错示例:rbindlist直接绑定向量报错 rbindlist(list(1:3, 4:6)) # Error in rbindlist(list(1:3, 4:6)) : # Item 1 of input is not a data.frame, data.table or list # 构造示例数据 temp_data <- structure(list(location_id = c(75L, 75L, 75L, 75L, 80L, 80L, 80L, 80L), date = structure(c(19144L, 19145L, 19146L, 19147L, 19144L, 19145L, 19146L, 19147L), class = c("IDate", "Date")), cases_cum = c(4289988, 4293027, 4295818, 4298654, 29570762, 29595892, 29621064, 29641606), population = c(8916185.49099959, 8916185.49099959, 8916185.49099959, 8916185.49099959, 66204314.9643178, 66204314.9643178, 66204314.9643178, 66204314.9643178), location_name = c("Austria", "Austria", "Austria", "Austria", "France", "France", "France", "France")), class = c("data.table", "data.frame"), row.names = c(NA, -8L), sorted = "location_id") # test1:rbindlist绑定报错 test1 <- rbindlist( lapply(split(temp_data, temp_data$location_id), function(x) { x <- x[order(x$date),] x$cases_daily <- c(NA,diff(x$cases_cum)) })) # Error in rbindlist(...) : Item 1 of input is not a data.frame, data.table or list # test2:do.call(rbind)仅返回cases_daily矩阵 test2 <- do.call(rbind, lapply(split(temp_data, temp_data$location_id), function(x) { x <- x[order(x$date)] x$cases_daily <- c(NA,diff(x$cases_cum)) }) ) # test2输出: # [,1] [,2] [,3] [,4] # 75 NA 3039 2791 2836 # 80 NA 25130 25172 20542 # test3:lapply结果仅为cases_daily向量 test3 <- lapply(split(temp_data, temp_data$location_id), function(x) { x <- x[order(x$date)] x$cases_daily <- c(NA,diff(x$cases_cum)) }) # test3输出: # $`75` # [1] NA 3039 2791 2836 # # $`80` # [1] NA 25130 25172 20542 # 手动处理可正常绑定 x <- temp_data[location_id == 80] x <- x[order(x$date)] x$cases_daily <- c(NA,diff(x$cases_cum)) y <- temp_data[location_id == 75] y <- y[order(y$date)] y$cases_daily <- c(NA,diff(y$cases_cum)) rbindlist(list(x,y)) # 输出包含所有列的完整data.table
原因分析
核心问题是**lapply中的匿名函数没有返回修改后的完整data.table**:
- R中函数的返回值是最后一行代码的结果,
x$cases_daily <- c(NA,diff(x$cases_cum))这行的返回值是赋值后的cases_daily向量,而非整个x - 因此
lapply返回的是每个分组的cases_daily向量列表,不是data.table列表,导致:rbindlist因输入不是数据框/表报错do.call(rbind)将向量拼接成矩阵
- 手动处理时,最后操作的是
x或y,默认返回整个对象,所以绑定正常
解决方案
方案1:修正lapply中的函数,显式返回完整data.table
在匿名函数末尾添加x(或return(x)),确保返回的是修改后的整个data.table:
test1_fixed <- rbindlist( lapply(split(temp_data, temp_data$location_id), function(x) { x <- x[order(x$date),] x$cases_daily <- c(NA,diff(x$cases_cum)) x # 显式返回完整data.table }))
方案2:使用data.table原生分组操作(更高效)
完全不需要split+lapply,直接用data.table的by参数实现分组计算,这是data.table的最优实践:
result <- temp_data[, cases_daily := c(NA, diff(cases_cum)), by = location_id][order(date),]
补充:rbindlist绑定向量的处理
如果需要用rbindlist绑定向量,需将向量转换为data.table或列表:
# 转换为data.table rbindlist(lapply(list(1:3, 4:6), function(v) data.table(value = v))) # 转换为列表 rbindlist(list(as.list(1:3), as.list(4:6)))
内容的提问来源于stack exchange,提问作者captaincaed
相关产品推荐
相关产品推荐

