如何将用户输入变量传入dplyr的group_by()与summarize()函数?
group_by() and summarize() Got it, let's tackle this problem! When you want to swap out hardcoded column names with user-provided variables in group_by() and summarize(), you need to account for dplyr's non-standard evaluation (NSE) behavior. Normally, dplyr expects bare column names, but when your column names are stored as variables (often strings from user input), you'll need tidy evaluation tools to make it work. Here are two straightforward approaches:
Approach 1: Using sym() and !! (Bang-Bang)
This method converts string variables into symbols that dplyr can recognize as column names:
library(dplyr) # Sample data frame (same as your example) df <- data.frame( 'Category' = c('a','c','a','a','b','a','b','b'), 'Amt' = c(100,300,200,400,500,1000,350,250), 'Flag' = c(0,1,1,1,0,1,1,0) ) # User-input variables (could come from a UI, console input, etc.) group_col <- "Category" # User's chosen grouping column value_col <- "Amt" # User's chosen value column for summation # Precompute totals (using the variable instead of hardcoding) rowCount <- nrow(df) totalAmt <- sum(df[[value_col]]) # Group using the variable g <- df %>% group_by(!!sym(group_col)) # !! unquotes the symbol created from the string # Summarize using the variable summ <- g %>% summarize( Count = n(), CountPercentage = n()*100/rowCount, TotalAmt = sum(!!sym(value_col)) ) # View the result summ
Approach 2: Using across() and all_of() (More Intuitive for Strings)
If you prefer a cleaner syntax for string inputs, across() paired with all_of() works great—it explicitly tells dplyr to look for column names matching your string variable:
library(dplyr) # Reuse the same data frame and user variables from above df <- data.frame( 'Category' = c('a','c','a','a','b','a','b','b'), 'Amt' = c(100,300,200,400,500,1000,350,250), 'Flag' = c(0,1,1,1,0,1,1,0) ) group_col <- "Category" value_col <- "Amt" rowCount <- nrow(df) totalAmt <- sum(df[[value_col]]) # Group with across() g <- df %>% group_by(across(all_of(group_col))) # Summarize with across() summ <- g %>% summarize( Count = n(), CountPercentage = n()*100/rowCount, TotalAmt = sum(across(all_of(value_col))) ) # View the result summ
What if the Variable is a Bare Name (Not a String)?
If your user input comes as a bare column name (e.g., group_col <- Category instead of a string), you can use the curly-curly operator {{}} directly:
group_col <- Category value_col <- Amt g <- df %>% group_by({{ group_col }}) summ <- g %>% summarize( Count = n(), CountPercentage = n()*100/rowCount, TotalAmt = sum({{ value_col }}) )
All these approaches will produce the exact same result as your original hardcoded code—just now using dynamic variables instead of fixed column names!
内容的提问来源于stack exchange,提问作者Suresh Subramaniam

