You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中为不同变量指定不同聚合函数(FUN)实现多变量分组聚合

Grouping Data with Different Aggregation Functions per Variable

Absolutely! You can easily apply different aggregation functions to separate variables when grouping your data in R. Let's use your sample dataset to demonstrate a few common, practical approaches:

First, let's recreate your sample data (so anyone following along can run the code):

dat <- data.frame(
  ID = rep(letters[1:3], each =3), 
  Plot = rep(1:3,3),
  Val1 = (1:9)*10, 
  Val2 = (1:9)*20
)

Here's what the dataset looks like:

ID Plot Val1 Val2
1  a    1   10   20
2  a    2   20   40
3  a    3   30   60
4  b    1   40   80
5  b    2   50  100
6  b    3   60  120
7  c    1   70  140
8  c    2   80  160
9  c    3   90  180

1. Using dplyr (Tidyverse Approach)

This is the most readable method for most R users, especially if you're working with the tidyverse ecosystem. Use group_by() to define your grouping variable(s), then summarize() to specify exactly which function to apply to each variable:

library(dplyr)

dat_summary <- dat %>%
  group_by(ID) %>%
  summarize(
    Total_Val1 = sum(Val1),  # Sum Val1 for each ID group
    Avg_Val2 = mean(Val2)   # Take the mean of Val2 for each ID group
  )

# View the result
dat_summary

Output:

# A tibble: 3 × 3
  ID    Total_Val1 Avg_Val2
  <chr>      <dbl>    <dbl>
1 a             60       40
2 b            150      100
3 c            240      160

2. Using data.table (Fast for Large Datasets)

If you're working with big datasets and need speed, data.table is the way to go. The syntax is concise and efficient:

library(data.table)

# Convert data.frame to data.table
setDT(dat)

dat_summary <- dat[, .(
  Total_Val1 = sum(Val1),
  Avg_Val2 = mean(Val2)
), by = ID]

# View the result
dat_summary

Output:

ID Total_Val1 Avg_Val2
1:  a         60       40
2:  b        150      100
3:  c        240      160

3. Base R Approach with aggregate()

You can also do this with base R's aggregate() function. Note that you'll need to pass a list of functions if you want different behavior per variable:

# For this method, we'll use a named list to map variables to functions
dat_summary <- aggregate(
  x = list(Val1 = dat$Val1, Val2 = dat$Val2),
  by = list(ID = dat$ID),
  FUN = function(x) c(sum = sum(x), mean = mean(x))
)

# Unnest the result for a cleaner table
dat_summary <- do.call(data.frame, dat_summary)
dat_summary

Output:

ID Val1.sum Val1.mean Val2.sum Val2.mean
1  a       60        20      120        40
2  b      150        50      300       100
3  c      240        80      480       160

(This returns both sum and mean for each variable, but you can adjust the function to only return what you need.)


内容的提问来源于stack exchange,提问作者theforestecologist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:53:05