You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Index列值筛选列表中DataFrame的指定列

问题描述

有一个包含多个列名相同的DataFrame的列表,需要根据每个DataFrame中Index列的取值,筛选出列名以该值结尾的列,同时保留Index列。尝试代码时出现object 'Index' not found错误。

示例输入

[[0]]
    a_1 b_1   c_1  a_2 b_2      c_2    Index
    3   red   no   2   yellow   yes    1

[[1]]
    a_1 b_1   c_1  a_2 b_2      c_2    Index
    3   red   no   2   yellow   yes    2

期望输出

[[0]]
    a_1 b_1   c_1     Index
    3   red   no       1

[[1]]
    a_2 b_2      c_2   Index
    2   yellow   yes   2

尝试代码及报错

尝试代码:

newlist<-lapply(samplelist,function(x) dplyr::select(ends_with(Index))) 

报错信息:

Error in is_character(match) : object 'Index' not found

补充数据(dput生成)

samplelist <- list(
  `1` = structure(list(ID = 12345, Com = structure(8296, class = "Date"), 
                       NCom = structure(8533, class = "Date"), 
                       a_1 = "Yes", b_1 = 160, c_1 = 160, d_1 = "No", 
                       e_1 = 0, f_1 = "No", g_1 = 0, h_1 = "Yes", 
                       a_2 = "Yes", b_2 = 155, c_2 = 155, d_2 = "No", 
                       e_2 = 0, d_2 = "No", e_2 = 0, f_2 = "Yes", 
                       Index = "1", Index_date = structure(9265, class = "Date")), row.names = 1L, class = "data.frame"),
  `2` = structure(list(Patient_ID = 22222, Com = structure(8296, class = "Date"), 
                       NCom = structure(8533, class = "Date"), 
                       a_1 = "Yes", b_1 = 160, c_1 = 160, d_1 = "No", 
                       e_1 = 0, f_1 = "No", g_1 = 0, h_1 = "Yes", 
                       a_2 = "Yes", b_2 = 155, c_2 = 155, d_2 = "No", 
                       e_2 = 0, d_2 = "No", e_2 = 0, f_2 = "Yes", 
                       Index = "2", Index_date = structure(8835, class = "Date")), row.names = 2L, class = "data.frame")
)

错误原因
  • 原代码未给dplyr::select指定数据源x,导致无法定位Index列
  • ends_with(Index)中的Index未从当前DataFrame中提取,上下文无法识别该变量

解决方案

方案1:使用dplyr结合lapply

修正代码,明确数据源并提取当前DataFrame的Index值:

library(dplyr)

newlist <- lapply(samplelist, function(x) {
  idx <- x$Index[[1]] # 提取当前DF的Index值(因每个DF仅一行)
  x %>%
    select(ends_with(idx), Index) # 筛选目标列+保留Index列
})

方案2:Base R实现(无需依赖dplyr)

用正则匹配筛选列:

newlist <- lapply(samplelist, function(x) {
  idx <- x$Index[[1]]
  # 匹配列名以"_idx"结尾 或 列名为Index的列
  keep_cols <- grepl(paste0("_", idx, "$"), colnames(x)) | colnames(x) == "Index"
  x[, keep_cols, drop = FALSE] # drop=FALSE确保结果为DataFrame格式
})

验证结果

运行上述任一代码后,每个DataFrame会仅保留以对应Index值结尾的列和Index列,与期望输出一致。

内容的提问来源于stack exchange,提问作者Hong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 19:25:20