如何对匹配正则的未知列执行条件mutate?嵌套data.frame处理
Hey there! Let's work through your problem to get the exact output you need, and clear up the across() vs vars() question along the way.
Step 1: Core Logic for Priority Filtering
First, we need a way to identify the highest-priority text in each target column, then keep only that text (removing lower-priority entries as NA). We'll use a priority vector to define the hierarchy, then build a helper function to apply this logic to any column.
Helper Function & Single DF Test
Here's how to process your sample df first, to verify the logic works:
library(tidyverse) # Your original sample data df <- tibble(ID_Value = 1:8, xyz001 = c("text4", NA, NA, NA, NA, NA, NA, "text2"), xyz002 = c(NA, NA, NA, "text3", "text1", NA, NA, NA), xyz003 = c(NA, "text1", NA, NA, "text2", NA, "text2", NA)) # Define priority order (highest to lowest) priority_levels <- c("text4", "text3", "text2", "text1") # Function to keep only highest-priority text in a column keep_highest_priority <- function(x) { # Skip processing if column is all NA non_na_vals <- x[!is.na(x)] if (length(non_na_vals) == 0) return(x) # Find the highest priority present in the column priority_indices <- match(non_na_vals, priority_levels) highest_priority_idx <- min(priority_indices, na.rm = TRUE) highest_val <- priority_levels[highest_priority_idx] # Keep only the highest-priority value, set others to NA ifelse(x == highest_val, x, NA) } # Apply to columns with numbers in their names processed_df <- df %>% mutate(across(matches("\\d"), keep_highest_priority)) # Check the result (matches your expected output!) processed_df
Step 2: Apply to Nested Data Frame
Now we can extend this to your df_nested using purrr::map() to process each nested data frame:
# Example nested data frame (adjust to match your actual structure) df_nested <- tibble(data = list(df, df)) # Process each nested df df_nested_processed <- df_nested %>% mutate(data = map(data, ~ .x %>% mutate(across(matches("\\d"), keep_highest_priority)))) # Inspect the first nested result df_nested_processed$data[[1]]
Step 3: across() vs vars()
Great question! Here's the breakdown:
vars()is a relic from older versions of dplyr (pre-1.0.0), used with functions likemutate_at()to select columns.across()is the modern, recommended replacement. It's more flexible: you can use all the standard select helpers (likematches(),starts_with()) directly inside it, combine multiple column selections in onemutate()call, and it integrates cleaner with the rest of the tidyverse syntax.
You should absolutely use across() here—vars() is no longer necessary for new code.
内容的提问来源于stack exchange,提问作者Don el Fuego Integer

