You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将用户输入变量传入dplyr的group_by()与summarize()函数?

How to Use User-Input Variables in dplyr's group_by() and summarize()

Got it, let's tackle this problem! When you want to swap out hardcoded column names with user-provided variables in group_by() and summarize(), you need to account for dplyr's non-standard evaluation (NSE) behavior. Normally, dplyr expects bare column names, but when your column names are stored as variables (often strings from user input), you'll need tidy evaluation tools to make it work. Here are two straightforward approaches:

Approach 1: Using sym() and !! (Bang-Bang)

This method converts string variables into symbols that dplyr can recognize as column names:

library(dplyr)

# Sample data frame (same as your example)
df <- data.frame(
  'Category' = c('a','c','a','a','b','a','b','b'), 
  'Amt' = c(100,300,200,400,500,1000,350,250), 
  'Flag' = c(0,1,1,1,0,1,1,0)
)

# User-input variables (could come from a UI, console input, etc.)
group_col <- "Category"  # User's chosen grouping column
value_col <- "Amt"       # User's chosen value column for summation

# Precompute totals (using the variable instead of hardcoding)
rowCount <- nrow(df)
totalAmt <- sum(df[[value_col]])

# Group using the variable
g <- df %>% 
  group_by(!!sym(group_col))  # !! unquotes the symbol created from the string

# Summarize using the variable
summ <- g %>% 
  summarize(
    Count = n(),
    CountPercentage = n()*100/rowCount,
    TotalAmt = sum(!!sym(value_col))
  )

# View the result
summ

Approach 2: Using across() and all_of() (More Intuitive for Strings)

If you prefer a cleaner syntax for string inputs, across() paired with all_of() works great—it explicitly tells dplyr to look for column names matching your string variable:

library(dplyr)

# Reuse the same data frame and user variables from above
df <- data.frame(
  'Category' = c('a','c','a','a','b','a','b','b'), 
  'Amt' = c(100,300,200,400,500,1000,350,250), 
  'Flag' = c(0,1,1,1,0,1,1,0)
)
group_col <- "Category"
value_col <- "Amt"

rowCount <- nrow(df)
totalAmt <- sum(df[[value_col]])

# Group with across()
g <- df %>% 
  group_by(across(all_of(group_col)))

# Summarize with across()
summ <- g %>% 
  summarize(
    Count = n(),
    CountPercentage = n()*100/rowCount,
    TotalAmt = sum(across(all_of(value_col)))
  )

# View the result
summ

What if the Variable is a Bare Name (Not a String)?

If your user input comes as a bare column name (e.g., group_col <- Category instead of a string), you can use the curly-curly operator {{}} directly:

group_col <- Category
value_col <- Amt

g <- df %>% 
  group_by({{ group_col }})

summ <- g %>% 
  summarize(
    Count = n(),
    CountPercentage = n()*100/rowCount,
    TotalAmt = sum({{ value_col }})
  )

All these approaches will produce the exact same result as your original hardcoded code—just now using dynamic variables instead of fixed column names!

内容的提问来源于stack exchange,提问作者Suresh Subramaniam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:26:13