如何使用purrr按类型对列表中的数据框基于id列执行全连接并输出列表格式
Grouped Full Join of Data Frames by Type (Using purrr)
Here's a straightforward solution that leverages purrr for iteration and dplyr for joining, which aligns with your preference for using the purrr package:
Step 1: Load Required Libraries
library(purrr) library(dplyr) library(stringr) # For string matching
Step 2: Define Your Input Data
list_example <- list(type1_a_b = data.frame(id = 1:3, a = 1:3, b = 4:6), type1_c_d = data.frame(id = 1:5, c = 1:5, d = 5:9), type2_e_f = data.frame(id = c(1,3,4), e = 1:3, f = 4:6), type2_g_h = data.frame(id = c(2,3,4), g = 1:3, h = 5:7)) data_types <- c("type1", "type2")
Step 3: Execute Grouped Full Join
result <- map(data_types, function(current_type) { # Filter list items whose names start with the current type list_example %>% keep(names(.) %>% str_detect(paste0("^", current_type))) %>% # Perform full join on all filtered data frames, using 'id' as the key reduce(full_join, by = "id") }) %>% # Assign names to the result list matching your data_types vector set_names(data_types)
Step 4: Verify the Output
When you print result, you'll get exactly the structured output you requested:
result #> $type1 #> id a b c d #> 1 1 1 4 1 5 #> 2 2 2 5 2 6 #> 3 3 3 6 3 7 #> 4 4 NA NA 4 8 #> 5 5 NA NA 5 9 #> #> $type2 #> id e f g h #> 1 1 1 4 NA NA #> 2 2 NA NA 1 5 #> 3 3 2 5 2 6 #> 4 4 3 6 3 7
How It Works
map()iterates over each value indata_types(i.e., "type1" and "type2"), running the inner logic for each group.keep()+str_detect()filters yourlist_exampleto only retain data frames whose names start with the current type (the regex^current_typeensures we match prefixes correctly).reduce(full_join, by = "id")takes the filtered sub-list of data frames and repeatedly appliesfull_jointo combine them all into one data frame, using "id" as the join key.set_names()ensures the final list has names matching yourdata_typesvector, making it easy to reference each grouped result for later processing.
内容的提问来源于stack exchange,提问作者Polina B
相关产品推荐
相关产品推荐

