You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对匹配正则的未知列执行条件mutate?嵌套data.frame处理

Solution for Priority-Based Text Filtering in Nested Data Frames

Hey there! Let's work through your problem to get the exact output you need, and clear up the across() vs vars() question along the way.

Step 1: Core Logic for Priority Filtering

First, we need a way to identify the highest-priority text in each target column, then keep only that text (removing lower-priority entries as NA). We'll use a priority vector to define the hierarchy, then build a helper function to apply this logic to any column.

Helper Function & Single DF Test

Here's how to process your sample df first, to verify the logic works:

library(tidyverse)

# Your original sample data
df <- tibble(ID_Value = 1:8, 
             xyz001 = c("text4", NA, NA, NA, NA, NA, NA, "text2"), 
             xyz002 = c(NA, NA, NA, "text3", "text1", NA, NA, NA), 
             xyz003 = c(NA, "text1", NA, NA, "text2", NA, "text2", NA))

# Define priority order (highest to lowest)
priority_levels <- c("text4", "text3", "text2", "text1")

# Function to keep only highest-priority text in a column
keep_highest_priority <- function(x) {
  # Skip processing if column is all NA
  non_na_vals <- x[!is.na(x)]
  if (length(non_na_vals) == 0) return(x)
  
  # Find the highest priority present in the column
  priority_indices <- match(non_na_vals, priority_levels)
  highest_priority_idx <- min(priority_indices, na.rm = TRUE)
  highest_val <- priority_levels[highest_priority_idx]
  
  # Keep only the highest-priority value, set others to NA
  ifelse(x == highest_val, x, NA)
}

# Apply to columns with numbers in their names
processed_df <- df %>%
  mutate(across(matches("\\d"), keep_highest_priority))

# Check the result (matches your expected output!)
processed_df

Step 2: Apply to Nested Data Frame

Now we can extend this to your df_nested using purrr::map() to process each nested data frame:

# Example nested data frame (adjust to match your actual structure)
df_nested <- tibble(data = list(df, df))

# Process each nested df
df_nested_processed <- df_nested %>%
  mutate(data = map(data, ~ .x %>%
                      mutate(across(matches("\\d"), keep_highest_priority))))

# Inspect the first nested result
df_nested_processed$data[[1]]

Step 3: across() vs vars()

Great question! Here's the breakdown:

  • vars() is a relic from older versions of dplyr (pre-1.0.0), used with functions like mutate_at() to select columns.
  • across() is the modern, recommended replacement. It's more flexible: you can use all the standard select helpers (like matches(), starts_with()) directly inside it, combine multiple column selections in one mutate() call, and it integrates cleaner with the rest of the tidyverse syntax.

You should absolutely use across() here—vars() is no longer necessary for new code.


内容的提问来源于stack exchange,提问作者Don el Fuego Integer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 20:42:30