如何以符合dplyr风格的DRY方式实现分组均值计算?
Great question! It’s super common to run into repetitive summarize() code when working with multiple variables in dplyr, especially when you’ve got lots of columns to aggregate. Luckily, dplyr has idiomatic tools built specifically for this that fit the DRY principle perfectly.
Recommended Approach: Use across() (dplyr 1.0.0+)
The across() function is the modern, preferred way to apply the same operation to multiple columns in dplyr. It’s flexible, readable, and eliminates all that repetitive mean(X) code.
Example Code
Here’s how you can rewrite your grouped mean calculation with across():
library(dplyr) # Generate test data first set.seed(42) df <- data.frame(GR=sample(1:2, 100, replace=TRUE), as.data.frame(replicate(20, rnorm(100)))) names(df)[-1] <- LETTERS[1:20] # DRY dplyr grouped means df %>% group_by(GR) %>% summarize(across(A:T, mean, .names = "mean.{col}"))
Breakdown:
across(A:T): Selects all columns fromAtoT(your 20 numeric variables). If you want to target all non-group columns without specifying a range, useacross(-GR)instead—this is even more flexible if your column list changes later.mean: The aggregation function we want to apply to each selected column..names = "mean.{col}": Defines the output column names.{col}is a placeholder that gets replaced with the original column name, so you’ll getmean.A,mean.B, ...,mean.Tjust like your original repetitive code.
Alternative (Deprecated): summarize_all()
Before across() was introduced, summarize_all() was the go-to for this use case. While it still works, it’s now marked as deprecated in favor of across(), so it’s better to use the modern approach. For reference, here’s what that would look like:
df %>% group_by(GR) %>% summarize_all(mean) %>% rename_with(~paste0("mean.", .), -GR)
This gives the same result, but requires an extra rename_with() step to get the mean.{col} column names, which makes across() the cleaner choice.
Verifying Consistency
The output from the across() method will match the structure of your original repetitive dplyr code, and the numeric values will be identical to what you get from aggregate(df[,-1], df[1], mean)—the only difference is the column naming (which we’ve intentionally matched to your original dplyr output).
This approach scales perfectly even if you have 50 or 100 columns to aggregate—no more copying and pasting mean(X) lines!
内容的提问来源于stack exchange,提问作者lebatsnok

