如何在嵌套列表中识别dataframe元素?(仅Base R实现)
用Base R定位嵌套列表中的data.frame元素
先定义测试用的嵌套列表:
test <- list( a = data.frame(x = 1), b = "foo", c = list( d = 1:5, e = data.frame(y = 1), f = "a", list(g = "hello") ) ) test #> $a #> x #> 1 1 #> #> $b #> [1] "foo" #> #> $c #> $c$d #> [1] 1 2 3 4 5 #> #> $c$e #> y #> 1 1 #> #> $c$f #> [1] "a" #> #> $c[[4]] #> $c[[4]]$g #> [1] "hello"
如果要找出列表中字符元素的位置,用rapply()就能直接实现——它会遍历展开所有元素,返回标记TRUE/FALSE的命名向量:
rapply(test, is.character) #> a.x b c.d c.e.y c.f c.g #> FALSE TRUE FALSE FALSE TRUE TRUE
但用rapply()查找data.frame时,它会把data.frame拆解成内部的列元素(比如结果里第一个标识是a.x而非列表层级的a),导致无法正确识别列表中的data.frame对象:
rapply(test, is.data.frame) #> a.x b c.d c.e.y c.f c.g #> FALSE FALSE FALSE FALSE FALSE FALSE
解决方法:自定义递归函数
可以用Base R写一个递归函数,遍历列表的每一层,直接判断每个元素是否为data.frame,保留原列表的层级结构:
find_dataframes <- function(x) { if (is.data.frame(x)) { return(TRUE) } if (is.list(x)) { return(lapply(x, find_dataframes)) } FALSE } # 调用函数 find_dataframes(test) #> $a #> [1] TRUE #> #> $b #> [1] FALSE #> #> $c #> $c$d #> [1] FALSE #> #> $c$e #> [1] TRUE #> #> $c$f #> [1] FALSE #> #> $c[[4]] #> $c[[4]]$g #> [1] FALSE
如果需要返回带路径的命名向量,可调整函数收集元素路径:
find_dataframes_named <- function(x, path = NULL) { result <- list() for (i in seq_along(x)) { # 处理命名元素和无命名元素的路径 elem_name <- names(x)[i] current_path <- if (is.null(path)) { if (is.na(elem_name)) paste0("[[", i, "]]") else elem_name } else { if (is.na(elem_name)) paste0(path, "[[", i, "]]") else paste(path, elem_name, sep = ".") } elem <- x[[i]] if (is.data.frame(elem)) { result[[current_path]] <- TRUE } else if (is.list(elem)) { result <- c(result, find_dataframes_named(elem, current_path)) } else { result[[current_path]] <- FALSE } } unlist(result) } # 调用函数 find_dataframes_named(test) #> a b c.d c.e c.f c[[4]]$g #> TRUE FALSE FALSE TRUE FALSE FALSE
这种递归方式会逐层遍历列表:遇到data.frame直接标记TRUE;遇到子列表就继续递归;其他类型标记FALSE,完全符合需求。
内容的提问来源于stack exchange,提问作者bretauv
相关产品推荐
相关产品推荐

