如何更简洁地将tibble的字符向量列按颜色家族拆分为多列?
Great question! Your initial approach using mutate() + map() + str_subset() totally gets the job done, but we can definitely make this cleaner and more maintainable. Let's break down the problem and explore two solid alternatives.
Why Your extract() Attempt Failed
You’re spot-on about the regex issue! tidyr::extract() works by pulling specific segments from a single string using capture groups, but it doesn’t know how to aggregate all matches for a category into one list. Your regex would only grab the first matching color for each family (e.g., "blue" for Cool, "red" for Warm) instead of collecting all relevant colors into each column. That’s why it didn’t produce the nested list columns you wanted.
Alternative 1: Use a Predefined Color Family Map (Most Concise & Maintainable)
This approach cuts down on repeated code by defining your color families once, then using purrr functions to map them across your data. It’s DRY (Don’t Repeat Yourself) and easy to update if your color categories change later.
library(tidyverse) # Step 1: Define your color families in a single list color_families <- list( Cool = c("blue", "green"), Warm = c("red", "pink", "yellow", "gold", "orange"), Neutral = c("ivory", "brown", "beige") ) # Your original tibble colours <- tribble( ~all, c('blue','green', 'red', 'pink', 'yellow', 'gold', 'orange', 'ivory', 'brown', 'beige'), c('green', 'red', 'pink', 'orange', 'ivory', 'beige') ) # Step 2: Map the families to create new columns colours %>% bind_cols( map_dfc(color_families, ~map(colours$all, intersect, .x)) )
How This Works:
map_dfc()iterates over each family incolor_familiesand returns a tibble where each column corresponds to a family.intersect(.x, colours$all)finds all colors in each row’salllist that belong to the current family, giving you the nested list columns you need.
Alternative 2: Unnest, Categorize, & Re-Nest (More Intuitive for Some)
If you prefer a more step-by-step approach that’s easier to follow visually, you can unnest the color lists, assign each color to a family, then re-nest and pivot back to wide format.
colours %>% # Add a row ID to keep track of original rows mutate(row_id = row_number()) %>% # Unnest the color lists into individual rows unnest(all) %>% # Assign each color to its family mutate(family = case_when( all %in% color_families$Cool ~ "Cool", all %in% color_families$Warm ~ "Warm", all %in% color_families$Neutral ~ "Neutral" )) %>% # Group by row and family, then re-nest colors group_by(row_id, family) %>% summarise(colors = list(all), .groups = "drop") %>% # Pivot back to wide format (one column per family) pivot_wider(names_from = family, values_from = colors) %>% # Join back with original data and clean up columns left_join(colours %>% mutate(row_id = row_number()), by = "row_id") %>% select(all, Cool, Warm, Neutral)
How This Works:
- Unnesting turns each color into its own row, making it easy to assign a family with
case_when(). - Re-nesting and pivoting wide puts the colors back into list columns organized by family, matching your desired output.
Final Output
Both methods will give you the exact result you wanted:
# A tibble: 2 × 4 all Cool Warm Neutral <list> <list> <list> <list> 1 <chr [10]> <chr [2]> <chr [5]> <chr [3]> 2 <chr [6]> <chr [1]> <chr [3]> <chr [2]>
内容的提问来源于stack exchange,提问作者user51462

