在R中为不同变量指定不同聚合函数(FUN)实现多变量分组聚合
Absolutely! You can easily apply different aggregation functions to separate variables when grouping your data in R. Let's use your sample dataset to demonstrate a few common, practical approaches:
First, let's recreate your sample data (so anyone following along can run the code):
dat <- data.frame( ID = rep(letters[1:3], each =3), Plot = rep(1:3,3), Val1 = (1:9)*10, Val2 = (1:9)*20 )
Here's what the dataset looks like:
ID Plot Val1 Val2 1 a 1 10 20 2 a 2 20 40 3 a 3 30 60 4 b 1 40 80 5 b 2 50 100 6 b 3 60 120 7 c 1 70 140 8 c 2 80 160 9 c 3 90 180
1. Using dplyr (Tidyverse Approach)
This is the most readable method for most R users, especially if you're working with the tidyverse ecosystem. Use group_by() to define your grouping variable(s), then summarize() to specify exactly which function to apply to each variable:
library(dplyr) dat_summary <- dat %>% group_by(ID) %>% summarize( Total_Val1 = sum(Val1), # Sum Val1 for each ID group Avg_Val2 = mean(Val2) # Take the mean of Val2 for each ID group ) # View the result dat_summary
Output:
# A tibble: 3 × 3 ID Total_Val1 Avg_Val2 <chr> <dbl> <dbl> 1 a 60 40 2 b 150 100 3 c 240 160
2. Using data.table (Fast for Large Datasets)
If you're working with big datasets and need speed, data.table is the way to go. The syntax is concise and efficient:
library(data.table) # Convert data.frame to data.table setDT(dat) dat_summary <- dat[, .( Total_Val1 = sum(Val1), Avg_Val2 = mean(Val2) ), by = ID] # View the result dat_summary
Output:
ID Total_Val1 Avg_Val2 1: a 60 40 2: b 150 100 3: c 240 160
3. Base R Approach with aggregate()
You can also do this with base R's aggregate() function. Note that you'll need to pass a list of functions if you want different behavior per variable:
# For this method, we'll use a named list to map variables to functions dat_summary <- aggregate( x = list(Val1 = dat$Val1, Val2 = dat$Val2), by = list(ID = dat$ID), FUN = function(x) c(sum = sum(x), mean = mean(x)) ) # Unnest the result for a cleaner table dat_summary <- do.call(data.frame, dat_summary) dat_summary
Output:
ID Val1.sum Val1.mean Val2.sum Val2.mean 1 a 60 20 120 40 2 b 150 50 300 100 3 c 240 80 480 160
(This returns both sum and mean for each variable, but you can adjust the function to only return what you need.)
内容的提问来源于stack exchange,提问作者theforestecologist

