You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tidyverse从嵌套列表提取列唯一值的优化实现问询

问题

已实现从嵌套列表(列表的列表)中提取指定列的唯一值,但希望采用更整洁的数据组织方式(避免当前的列表组织形式),并完全利用tidyverse工具(如purrr)替代for循环与lapply,实现相同输出。现有代码如下:

#Load libraries
library(tidyverse)
library(stringi)

#Generate datasets to create example
dat1 <- tibble(col1 = runif(10), col2 = stri_rand_strings(10, 5))
dat2 <- tibble(col1 = runif(10), col2 = stri_rand_strings(10, 5))
dat3 <- tibble(col1 = runif(10), col2 = stri_rand_strings(10, 5))
dat4 <- tibble(col1 = runif(10), col2 = stri_rand_strings(10, 5))

#Generate lists from the generated datasets (3 of them for example)
list1 <- list(id1 = dat1, id2 = dat2)
list2 <- list(id1 = dat3)
list3 <- list(id1 = dat4)
#Generate list of lists
final_list <- list(dataset1 = list1, dataset2 = list2, dataset3 = list3)

#Select unique cases from the list of lists
#First loop for each list (1, 2 and 3) and then within list select unique cases with lapply for col2

data <- NULL

for (i in final_list) {

single_cases <- bind_rows(lapply(i, function(x) x %>% select(col2) %>% distinct(col2)))

data <- rbind(data, single_cases)  
  
}
解决方案

可以利用purrr系列函数结合tidyverse管道语法,完全替代循环与lapply,同时输出整洁的tibble格式:

# 加载必要包(已加载可省略)
library(tidyverse)

# 简洁实现:展平嵌套列表并提取唯一值
data_clean <- final_list %>%
  # 对每一层子列表合并为tibble,再合并所有层
  map_dfr(bind_rows) %>%
  # 选择目标列
  select(col2) %>%
  # 提取唯一值
  distinct(col2)

另一种分层处理的直观方式

如果想明确处理嵌套结构的每一层,也可以用map_depth深入到第二层列表提取目标列,再展平合并:

data_clean <- final_list %>%
  # 深入到第二层列表,提取col2列
  map_depth(2, ~ select(.x, col2)) %>%
  # 展平所有嵌套结构为单个tibble
  flatten_dfr() %>%
  # 去重
  distinct(col2)

说明

  • 两种方式都完全遵循tidyverse的函数式编程风格,避免了手动循环和lapply
  • 输出结果为tibble,比原代码生成的data.frame更整洁,自带类型信息和友好的打印格式
  • 代码逻辑连贯易读:从嵌套列表 -> 合并为统一表格 -> 选列 -> 去重

内容的提问来源于stack exchange,提问作者J. Lan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 00:09:22