如何在R中使用索引列表分割或子集化嵌套dataframe?
解决嵌套DataFrame结合
createDataPartition的子集化问题 看起来你遇到的核心问题是没把partitionindex里的索引和nested_df中对应的子DataFrame正确匹配,直接用列表作为下标触发了类型错误。咱们一步步来搞定:
错误原因分析
你的partitionindex是个长度为4的嵌套列表——外层的4个元素刚好对应nested_df里的4个League(比如[[1]]对应F1,[[4]]对应D1),每个元素内部的$Resample1才是该League数据集的训练行索引。之前的代码既没把索引和子DataFrame一一绑定,还直接把列表当下标用,R自然会报错"invalid subscript type 'list'"。
解决方案:用purrr::map2一一对应处理
因为nested_df$data(4个子DataFrame)和partitionindex(4组索引)是严格一一对应的,我们可以用map2同时遍历这两个列表,对每个子DataFrame应用对应的索引:
library(tidyverse) library(caret) # 基于你的已有数据执行以下代码 nested_df <- nested_df %>% mutate( train = map2( .x = data, # 遍历每个League的子DataFrame .y = partitionindex, # 遍历对应League的索引组 .f = function(df, idx) df[idx$Resample1, ] # 提取Resample1的索引来子集化 ) )
代码细节解释
map2会自动把data里的子DataFrame和partitionindex里的对应索引组配对,再传递给自定义函数。- 自定义函数里的
idx$Resample1专门取出每组索引中的训练行号,用这个行号去切分对应的子DataFrame,完美规避了列表下标错误。
验证结果是否正确
可以简单检查生成的训练集是否符合预期:
# 查看第一个League的训练集行数 nrow(nested_df$train[[1]]) # 对比对应索引组的元素数量 length(partitionindex[[1]]$Resample1)
两者的数值应该完全一致,说明子集化成功。
内容的提问来源于stack exchange,提问作者Chewyham
相关产品推荐
相关产品推荐

