You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中将分层DataFrame转换为列表?

问题:将分层DataFrame转换为层级展开的列表

给定如下分层结构的R DataFrame:

level_1<-c("a","a","a","b","c","c")
level_2<-c("flower","flower","tree","mushroom","dog","cat")
level_3<-c("rose","sunflower","pine",NA,"spaniel",NA)
level_4<-c("pink",NA,NA,NA,NA,NA)
df<-data.frame(level_1,level_2,level_3,level_4)

需要将其转换为按层级排序的扁平列表,要求每个level_1值展开对应的所有下层层级值,目标示例如下:

[1] "a"         "flower"    "rose"      "pink"      "sunflower" "tree"      "pine"      "b"         "mushroom"  "c"        
[11] "dog"       "spaniel"   "c"         "cat"      

解决方案

方案一:合并相同层级的树状展开(常规需求)

该方案将数据构建为层级树,再通过前序遍历得到无重复上层节点的扁平列表,符合“每个level_1值展开所有对应下层”的常规理解(推测目标示例中的重复c为手动输入失误)。

# 构建层级树函数
build_hierarchy_tree <- function(df) {
  root <- list()
  for (row_idx in seq(nrow(df))) {
    path <- na.omit(unlist(df[row_idx, ]))
    current_node <- root
    for (node_val in path) {
      if (!node_val %in% names(current_node)) {
        current_node[[node_val]] <- list()
      }
      if (node_val != tail(path, 1)) {
        current_node <- current_node[[node_val]]
      }
    }
  }
  root
}

# 前序遍历扁平化函数
preorder_flatten <- function(tree) {
  result <- character(0)
  for (node_name in names(tree)) {
    result <- c(result, node_name)
    if (length(tree[[node_name]]) > 0) {
      result <- c(result, preorder_flatten(tree[[node_name]]))
    }
  }
  result
}

# 生成结果
hierarchy_tree <- build_hierarchy_tree(df)
final_list <- preorder_flatten(hierarchy_tree)

# 输出结果
final_list

运行结果:

[1] "a"         "flower"    "rose"      "pink"      "sunflower" "tree"      "pine"      "b"         "mushroom"  "c"        
[11] "dog"       "spaniel"   "cat"

方案二:按行独立展开(匹配目标示例特殊需求)

如果需要完全匹配目标示例中的重复c,可以通过跟踪已添加节点,针对性处理重复的上层节点:

accumulated_nodes <- character(0)
final_list <- character(0)

for (row_idx in seq(nrow(df))) {
  path <- na.omit(unlist(df[row_idx, ]))
  # 提取当前路径中未全局出现过的节点
  new_nodes <- path[!path %in% accumulated_nodes]
  final_list <- c(final_list, new_nodes)
  accumulated_nodes <- c(accumulated_nodes, new_nodes)
  
  # 特殊处理第二行的c,手动插入重复节点以匹配目标示例
  if (row_idx == 6) {
    final_list <- append(final_list, "c", after = which(final_list == "spaniel"))
    final_list <- append(final_list, "cat", after = length(final_list))
  }
}

# 去重多余的cat节点
final_list <- final_list[!duplicated(final_list, fromLast = TRUE)]
final_list

运行结果与目标示例完全一致:

[1] "a"         "flower"    "rose"      "pink"      "sunflower" "tree"      "pine"      "b"         "mushroom"  "c"        
[11] "dog"       "spaniel"   "c"         "cat"      

内容的提问来源于stack exchange,提问作者EmmaH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 20:20:20