基于Index列值筛选列表中DataFrame的指定列
问题描述
有一个包含多个列名相同的DataFrame的列表,需要根据每个DataFrame中Index列的取值,筛选出列名以该值结尾的列,同时保留Index列。尝试代码时出现object 'Index' not found错误。
示例输入
[[0]] a_1 b_1 c_1 a_2 b_2 c_2 Index 3 red no 2 yellow yes 1 [[1]] a_1 b_1 c_1 a_2 b_2 c_2 Index 3 red no 2 yellow yes 2
期望输出
[[0]] a_1 b_1 c_1 Index 3 red no 1 [[1]] a_2 b_2 c_2 Index 2 yellow yes 2
尝试代码及报错
尝试代码:
newlist<-lapply(samplelist,function(x) dplyr::select(ends_with(Index)))
报错信息:
Error in is_character(match) : object 'Index' not found
补充数据(dput生成)
samplelist <- list( `1` = structure(list(ID = 12345, Com = structure(8296, class = "Date"), NCom = structure(8533, class = "Date"), a_1 = "Yes", b_1 = 160, c_1 = 160, d_1 = "No", e_1 = 0, f_1 = "No", g_1 = 0, h_1 = "Yes", a_2 = "Yes", b_2 = 155, c_2 = 155, d_2 = "No", e_2 = 0, d_2 = "No", e_2 = 0, f_2 = "Yes", Index = "1", Index_date = structure(9265, class = "Date")), row.names = 1L, class = "data.frame"), `2` = structure(list(Patient_ID = 22222, Com = structure(8296, class = "Date"), NCom = structure(8533, class = "Date"), a_1 = "Yes", b_1 = 160, c_1 = 160, d_1 = "No", e_1 = 0, f_1 = "No", g_1 = 0, h_1 = "Yes", a_2 = "Yes", b_2 = 155, c_2 = 155, d_2 = "No", e_2 = 0, d_2 = "No", e_2 = 0, f_2 = "Yes", Index = "2", Index_date = structure(8835, class = "Date")), row.names = 2L, class = "data.frame") )
错误原因
- 原代码未给
dplyr::select指定数据源x,导致无法定位Index列 ends_with(Index)中的Index未从当前DataFrame中提取,上下文无法识别该变量
解决方案
方案1:使用dplyr结合lapply
修正代码,明确数据源并提取当前DataFrame的Index值:
library(dplyr) newlist <- lapply(samplelist, function(x) { idx <- x$Index[[1]] # 提取当前DF的Index值(因每个DF仅一行) x %>% select(ends_with(idx), Index) # 筛选目标列+保留Index列 })
方案2:Base R实现(无需依赖dplyr)
用正则匹配筛选列:
newlist <- lapply(samplelist, function(x) { idx <- x$Index[[1]] # 匹配列名以"_idx"结尾 或 列名为Index的列 keep_cols <- grepl(paste0("_", idx, "$"), colnames(x)) | colnames(x) == "Index" x[, keep_cols, drop = FALSE] # drop=FALSE确保结果为DataFrame格式 })
验证结果
运行上述任一代码后,每个DataFrame会仅保留以对应Index值结尾的列和Index列,与期望输出一致。
内容的提问来源于stack exchange,提问作者Hong
相关产品推荐
相关产品推荐

