You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中对数据框进行聚合与分组?附示例数据

R Data Aggregation & Grouping for Your Modified USArrests Dataset

Got it, let's walk through how to handle grouping and aggregation on your dataset using two common R approaches: base R and the tidyverse's dplyr package. First, let's make sure everyone can replicate your data setup:

# Load and prepare the sample data
df <- USArrests
df$ID <- c(1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 3, 3, 3, 3, 3, 3, 3, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 2, 2, 2, 2)
df$Year <- c(2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2012, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2011, 2015, 2012, 2013)

1. Base R Approach

Base R has the aggregate() function which works well for straightforward grouping tasks.

Group by a Single Variable (e.g., ID)

Calculate the mean of all arrest-related columns for each ID:

# Mean of numeric columns (excluding ID/Year, since they're grouping variables)
aggregate(cbind(Murder, Assault, UrbanPop, Rape) ~ ID, data = df, FUN = mean, na.rm = TRUE)

# Or sum a specific column, like total Assaults per ID
aggregate(Assault ~ ID, data = df, FUN = sum)

Group by Multiple Variables (ID + Year)

Count observations and compute median Rape rates for each ID-Year combination:

# Number of rows per ID-Year group
aggregate(Murder ~ ID + Year, data = df, FUN = length) # Using Murder as a placeholder column

# Median Rape rate per ID-Year group
aggregate(Rape ~ ID + Year, data = df, FUN = median, na.rm = TRUE)

2. dplyr/Tidyverse Approach

This is the more flexible and readable method for complex grouping tasks. First, load the package if you haven't already:

library(dplyr)

Group by ID and Compute Custom Aggregates

Use group_by() to define groups, then summarize() to calculate stats:

df %>%
  group_by(ID) %>%
  summarize(
    avg_murder = mean(Murder, na.rm = TRUE),
    total_assault = sum(Assault, na.rm = TRUE),
    num_records = n(), # Count rows in each group
    max_urbanpop = max(UrbanPop, na.rm = TRUE)
  ) %>%
  ungroup() # Always ungroup after operations to avoid unexpected behavior later

Group by ID + Year

Easily compute multiple stats across groups with across():

df %>%
  group_by(ID, Year) %>%
  summarize(
    across(c(Murder, Rape), mean, na.rm = TRUE), # Mean of Murder/Rape per group
    total_urbanpop = sum(UrbanPop, na.rm = TRUE)
  ) %>%
  ungroup()

Grouped Transformation (Add Stats to Original Data)

If you want to append group-level stats to every row (e.g., average Murder rate for each ID):

df_with_group_stats <- df %>%
  group_by(ID) %>%
  mutate(avg_murder_for_id = mean(Murder, na.rm = TRUE)) %>%
  ungroup()

# View the result
head(df_with_group_stats)

Bonus: Useful Aggregation Functions

Here are some common functions you can use in summarize() or aggregate():

  • n(): Count rows in a group
  • sum(), mean(), median(): Summarize numeric values
  • max(), min(): Get extreme values
  • sd(): Calculate standard deviation
  • first(), last(): Retrieve the first/last value in a group

内容的提问来源于stack exchange,提问作者WoeIs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:00:07