如何在rapply处理嵌套列表的输出中,将id信息整合到行名?
问题描述
在三级嵌套列表场景中,需定位至第三层(resource)检查数据是否包含特殊字符,当前已通过以下R代码实现基础检查:
obj1 <- list(resource = list(bodyPart = c("leg", "arm", "knee"),side = c("LEFT", "RIGHT", "LEFT"), device = c("LLI", "LSM", "GHT"), id = c("AA", "BB", "CC")) %>% as.data.frame(), cat = list(lab = c("aa", "bb", "cc"), id = c(23, 24, 25)) %>% as.data.frame()) obj2 <- list(resource = list(bodyPart = c("leg", "arm", "knee"), side = c("LEFT", "LEFT", "LEFT"), device = c("GOM", "LSM", "YYY"), id = c("ZZ", "DD", "FF")) %>% as.data.frame()) x <- list(foo = c(fer = "wdb", obj1), bar = obj2) library(tibble) data.frame(has_invalid_character = x |> rapply(f = \(node) grepl("[^\x01-\x7F]", node)), content = x |> rapply(f = \(node) node) )
现在需要将数据中id列的信息添加至输出数据框的行名,替换原行名中的数字后缀(例如将foo.resource.bodyPart1替换为foo.resource.bodyPart.AA),以便精准定位数据所在位置,期望输出示例如下:
has_invalid_character content foo.fer FALSE wdb foo.resource.bodyPart1.AA FALSE leg foo.resource.bodyPart2.BB FALSE arm foo.resource.bodyPart3.CC FALSE knee foo.resource.side1.AA FALSE LEFT foo.resource.side2.BB FALSE RIGHT foo.resource.side3.CC FALSE LEFT foo.cat.lab1.23 FALSE aa foo.cat.lab1.24 FALSE bb foo.cat.lab1.25 FALSE cc
解决方案
可以通过递归提取id列映射关系,再批量替换行名的方式实现需求,具体代码如下:
library(dplyr) library(tibble) # 原始数据定义 obj1 <- list( resource = list(bodyPart = c("leg", "arm", "knee"), side = c("LEFT", "RIGHT", "LEFT"), device = c("LLI", "LSM", "GHT"), id = c("AA", "BB", "CC")) %>% as.data.frame(), cat = list(lab = c("aa", "bb", "cc"), id = c(23, 24, 25)) %>% as.data.frame() ) obj2 <- list( resource = list(bodyPart = c("leg", "arm", "knee"), side = c("LEFT", "LEFT", "LEFT"), device = c("GOM", "LSM", "YYY"), id = c("ZZ", "DD", "FF")) %>% as.data.frame() ) x <- list(foo = c(fer = "wdb", obj1), bar = obj2) # 1. 递归提取所有含id列的数据框的路径与对应id值 extract_ids <- function(lst, parent_path = "") { id_list <- list() for (name in names(lst)) { current_path <- if (parent_path == "") name else paste(parent_path, name, sep = ".") element <- lst[[name]] if (is.data.frame(element) && "id" %in% colnames(element)) { id_list[[current_path]] <- element$id } else if (is.list(element)) { id_list <- c(id_list, extract_ids(element, current_path)) } } return(id_list) } id_mapping <- extract_ids(x) # 2. 生成原始检查结果 result_df <- data.frame( has_invalid_character = x |> rapply(f = \(node) grepl("[^\x01-\x7F]", node)), content = x |> rapply(f = \(node) node) ) # 3. 替换行名中的数字后缀为对应id值 new_rownames <- sapply(rownames(result_df), function(rn) { parts <- strsplit(rn, "\\.")[[1]] num_pos <- grep("\\d+$", parts) if (length(num_pos) > 0) { parent_path <- paste(parts[1:(num_pos-1)], collapse = ".") idx <- as.integer(sub("\\D", "", parts[num_pos])) if (parent_path %in% names(id_mapping)) { id_val <- id_mapping[[parent_path]][idx] parts[num_pos] <- sub("(\\D+)(\\d+)", paste0("\\1.", id_val), parts[num_pos]) } } paste(parts, collapse = ".") }) rownames(result_df) <- new_rownames # 输出结果 print(result_df)
代码说明
extract_ids函数:递归遍历嵌套列表,记录每个包含id列的数据框的完整路径,以及对应的id值列表,建立路径与id的映射关系。- 行名替换逻辑:拆分每个行名的路径,识别出带数字后缀的部分,通过映射关系找到对应位置的
id值,将数字后缀替换为.id值的格式,最终拼接成新行名。
内容的提问来源于stack exchange,提问作者Rara
相关产品推荐
相关产品推荐

