R语言中使用aggregate函数时如何为同一函数的两次调用传递不同参数
Great question! Let's break this down. First, let's confirm we're working with the mpg dataset from the ggplot2 package (I’ll assume that’s loaded for context).
First, your initial year processing code is solid:
library(ggplot2) data(mpg) year <- mpg$year year[mpg$year > 2006] <- NA
Now, to your core question: Is there a way to run aggregate() with two sum calls (one with na.rm=TRUE, one with na.rm=FALSE) without writing a custom function(x)?
Short Answer
Strictly speaking, no—if you want to pass different na.rm values to each sum call in a single aggregate() call, you can’t avoid some form of function wrapping (even if it’s implicit via plyr::each). But there are workarounds that are close to what you want, or alternative approaches that avoid custom functions entirely.
Option 1: Use plyr::each() with concise function wrappers (closest to your original idea)
Your initial attempt with plyr::each() was on the right track, but you can’t pass duplicate na.rm arguments directly. Instead, you can embed the na.rm parameter into each sum call inside each():
library(plyr) aggregate(year, by = list(model = mpg$model), FUN = each(sum_na_rm = function(x) sum(x, na.rm = TRUE), sum_no_na = function(x) sum(x, na.rm = FALSE)))
While this uses anonymous functions, it’s a concise way to get both sums in one call and aligns with the spirit of your original approach.
Option 2: Run two separate aggregate() calls and merge results (no custom functions at all)
If you want to avoid any custom functions entirely, the simplest approach is to compute each sum separately and combine the results:
# Calculate sum with NA removed sum_rm <- aggregate(year, by = list(model = mpg$model), FUN = sum, na.rm = TRUE) # Calculate sum with NA retained (returns NA if any value in the group is NA) sum_no_rm <- aggregate(year, by = list(model = mpg$model), FUN = sum, na.rm = FALSE) # Merge and clean up the results merged_results <- merge(sum_rm, sum_no_rm, by = "Group.1", suffixes = c("_na_rm", "_no_na")) names(merged_results)[1] <- "model" # Rename group column for clarity
This uses only base R, no custom functions, and gives you the two sum columns you need.
Why you can’t do it with a single aggregate() call without function wrappers
The aggregate() function passes the same arguments to every function in FUN—so if you pass na.rm=TRUE, both sum calls would use that. There’s no built-in way to pass different parameters to multiple instances of the same function in a single aggregate() call without wrapping each function to set its own parameters.
内容的提问来源于stack exchange,提问作者pgitti

