You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在rapply处理嵌套列表的输出中,将id信息整合到行名?

问题描述

在三级嵌套列表场景中,需定位至第三层(resource)检查数据是否包含特殊字符,当前已通过以下R代码实现基础检查:

obj1 <- list(resource = list(bodyPart = c("leg", "arm", "knee"),side = c("LEFT", "RIGHT", "LEFT"), device = c("LLI", "LSM", "GHT"), id = c("AA", "BB", "CC")) %>%  as.data.frame(), cat = list(lab = c("aa", "bb", "cc"), id = c(23, 24, 25)) %>%  as.data.frame())
obj2 <- list(resource = list(bodyPart = c("leg", "arm", "knee"), side = c("LEFT", "LEFT", "LEFT"), device = c("GOM", "LSM", "YYY"), id = c("ZZ", "DD", "FF")) %>%  as.data.frame())

x <- list(foo = c(fer = "wdb", obj1), bar = obj2)

library(tibble)

data.frame(has_invalid_character = x |> rapply(f = \(node) grepl("[^\x01-\x7F]", node)),
           content =  x |> rapply(f = \(node) node)
           ) 

现在需要将数据中id列的信息添加至输出数据框的行名,替换原行名中的数字后缀(例如将foo.resource.bodyPart1替换为foo.resource.bodyPart.AA),以便精准定位数据所在位置,期望输出示例如下:

has_invalid_character content
foo.fer                                   FALSE     wdb
foo.resource.bodyPart1.AA                 FALSE     leg
foo.resource.bodyPart2.BB                 FALSE     arm
foo.resource.bodyPart3.CC                 FALSE    knee
foo.resource.side1.AA                     FALSE    LEFT
foo.resource.side2.BB                     FALSE   RIGHT
foo.resource.side3.CC                     FALSE    LEFT
foo.cat.lab1.23                           FALSE     aa
foo.cat.lab1.24                           FALSE     bb
foo.cat.lab1.25                           FALSE     cc
解决方案

可以通过递归提取id列映射关系,再批量替换行名的方式实现需求,具体代码如下:

library(dplyr)
library(tibble)

# 原始数据定义
obj1 <- list(
  resource = list(bodyPart = c("leg", "arm", "knee"),
                  side = c("LEFT", "RIGHT", "LEFT"), 
                  device = c("LLI", "LSM", "GHT"), 
                  id = c("AA", "BB", "CC")) %>% as.data.frame(),
  cat = list(lab = c("aa", "bb", "cc"), 
             id = c(23, 24, 25)) %>% as.data.frame()
)
obj2 <- list(
  resource = list(bodyPart = c("leg", "arm", "knee"), 
                  side = c("LEFT", "LEFT", "LEFT"), 
                  device = c("GOM", "LSM", "YYY"), 
                  id = c("ZZ", "DD", "FF")) %>% as.data.frame()
)

x <- list(foo = c(fer = "wdb", obj1), bar = obj2)

# 1. 递归提取所有含id列的数据框的路径与对应id值
extract_ids <- function(lst, parent_path = "") {
  id_list <- list()
  for (name in names(lst)) {
    current_path <- if (parent_path == "") name else paste(parent_path, name, sep = ".")
    element <- lst[[name]]
    if (is.data.frame(element) && "id" %in% colnames(element)) {
      id_list[[current_path]] <- element$id
    } else if (is.list(element)) {
      id_list <- c(id_list, extract_ids(element, current_path))
    }
  }
  return(id_list)
}

id_mapping <- extract_ids(x)

# 2. 生成原始检查结果
result_df <- data.frame(
  has_invalid_character = x |> rapply(f = \(node) grepl("[^\x01-\x7F]", node)),
  content =  x |> rapply(f = \(node) node)
)

# 3. 替换行名中的数字后缀为对应id值
new_rownames <- sapply(rownames(result_df), function(rn) {
  parts <- strsplit(rn, "\\.")[[1]]
  num_pos <- grep("\\d+$", parts)
  if (length(num_pos) > 0) {
    parent_path <- paste(parts[1:(num_pos-1)], collapse = ".")
    idx <- as.integer(sub("\\D", "", parts[num_pos]))
    if (parent_path %in% names(id_mapping)) {
      id_val <- id_mapping[[parent_path]][idx]
      parts[num_pos] <- sub("(\\D+)(\\d+)", paste0("\\1.", id_val), parts[num_pos])
    }
  }
  paste(parts, collapse = ".")
})

rownames(result_df) <- new_rownames

# 输出结果
print(result_df)

代码说明

  • extract_ids函数:递归遍历嵌套列表,记录每个包含id列的数据框的完整路径,以及对应的id值列表,建立路径与id的映射关系。
  • 行名替换逻辑:拆分每个行名的路径,识别出带数字后缀的部分,通过映射关系找到对应位置的id值,将数字后缀替换为.id值的格式,最终拼接成新行名。

内容的提问来源于stack exchange,提问作者Rara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 18:20:14