使用tidyverse从嵌套列表提取列唯一值的优化实现问询
问题
已实现从嵌套列表(列表的列表)中提取指定列的唯一值,但希望采用更整洁的数据组织方式(避免当前的列表组织形式),并完全利用tidyverse工具(如purrr)替代for循环与lapply,实现相同输出。现有代码如下:
#Load libraries library(tidyverse) library(stringi) #Generate datasets to create example dat1 <- tibble(col1 = runif(10), col2 = stri_rand_strings(10, 5)) dat2 <- tibble(col1 = runif(10), col2 = stri_rand_strings(10, 5)) dat3 <- tibble(col1 = runif(10), col2 = stri_rand_strings(10, 5)) dat4 <- tibble(col1 = runif(10), col2 = stri_rand_strings(10, 5)) #Generate lists from the generated datasets (3 of them for example) list1 <- list(id1 = dat1, id2 = dat2) list2 <- list(id1 = dat3) list3 <- list(id1 = dat4) #Generate list of lists final_list <- list(dataset1 = list1, dataset2 = list2, dataset3 = list3) #Select unique cases from the list of lists #First loop for each list (1, 2 and 3) and then within list select unique cases with lapply for col2 data <- NULL for (i in final_list) { single_cases <- bind_rows(lapply(i, function(x) x %>% select(col2) %>% distinct(col2))) data <- rbind(data, single_cases) }
解决方案
可以利用purrr系列函数结合tidyverse管道语法,完全替代循环与lapply,同时输出整洁的tibble格式:
# 加载必要包(已加载可省略) library(tidyverse) # 简洁实现:展平嵌套列表并提取唯一值 data_clean <- final_list %>% # 对每一层子列表合并为tibble,再合并所有层 map_dfr(bind_rows) %>% # 选择目标列 select(col2) %>% # 提取唯一值 distinct(col2)
另一种分层处理的直观方式
如果想明确处理嵌套结构的每一层,也可以用map_depth深入到第二层列表提取目标列,再展平合并:
data_clean <- final_list %>% # 深入到第二层列表,提取col2列 map_depth(2, ~ select(.x, col2)) %>% # 展平所有嵌套结构为单个tibble flatten_dfr() %>% # 去重 distinct(col2)
说明
- 两种方式都完全遵循tidyverse的函数式编程风格,避免了手动循环和lapply
- 输出结果为
tibble,比原代码生成的data.frame更整洁,自带类型信息和友好的打印格式 - 代码逻辑连贯易读:从嵌套列表 -> 合并为统一表格 -> 选列 -> 去重
内容的提问来源于stack exchange,提问作者J. Lan
相关产品推荐
相关产品推荐

