R语言中嵌套列表转DataFrame时重复列名规范化处理问题——基于过往相关方案的进阶需求
The root issue here is that as_tibble()'s default "unique" name repair appends ...N to duplicate column names, which doesn't match your desired matches.id, matches.id.1, matches.id.2 pattern. Instead of fixing column names after the fact, you can replace the default name repair with a custom function that generates clean, sequential suffixes directly in your pipeline.
Step 1: Define a Custom Name Repair Function
First, create a function that handles duplicate column names by appending .1, .2, etc., while keeping the first occurrence of each name unchanged:
custom_name_repair <- function(col_names) { # Group names and add sequential suffixes to duplicates ave(col_names, col_names, FUN = function(x) { if (length(x) == 1) { x # Keep unique names as-is } else { # Keep first name unchanged, append .1, .2 to subsequent duplicates c(x[1], paste0(x[-1], ".", seq_along(x[-1]))) } }) }
Step 2: Integrate the Function into Your Pipeline
Replace .name_repair = "unique" in your as_tibble() call with this custom function. Here's your updated pipeline:
output_df <- raw_list %>% rrapply(f = function(x) replace(x, is.null(x), NA)) %>% purrr::map(unlist) %>% purrr::map(t) %>% {suppressMessages(purrr::map(., as_tibble, .name_repair = custom_name_repair))} %>% dplyr::bind_rows() %>% {suppressMessages(readr::type_convert(.))} %>% as.data.frame()
How This Works
- When converting each transposed unlisted list to a tibble,
custom_name_repairchecks for duplicate column names. - For columns that appear multiple times, it retains the original name for the first occurrence and appends
.1,.2, etc., to subsequent duplicates—exactly the pattern you want (e.g.,matches.id,matches.id.1,matches.id.2). - This fixes the column naming issue at the source, eliminating the need for post-processing with
tidyr::unite()andseparate().
Example Test Case
If you test with a sample nested list like this:
raw_list <- list( list(user_id = 1, matches.id = 101, score = 85, matches.id = 102), list(user_id = 2, matches.id = 201, matches.id = 202, matches.id = 203) )
Your output DataFrame will have clean column names: user_id, matches.id, score, matches.id.1, matches.id.2—no messy ...N suffixes.
内容的提问来源于stack exchange,提问作者Canovice

