如何在R语言中对数据框进行聚合与分组?附示例数据
Got it, let's walk through how to handle grouping and aggregation on your dataset using two common R approaches: base R and the tidyverse's dplyr package. First, let's make sure everyone can replicate your data setup:
# Load and prepare the sample data df <- USArrests df$ID <- c(1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 3, 3, 3, 3, 3, 3, 3, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 2, 2, 2, 2) df$Year <- c(2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2012, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2012, 2013, 2017, 2017, 2017, 2015, 2011, 2015, 2012, 2013)
1. Base R Approach
Base R has the aggregate() function which works well for straightforward grouping tasks.
Group by a Single Variable (e.g., ID)
Calculate the mean of all arrest-related columns for each ID:
# Mean of numeric columns (excluding ID/Year, since they're grouping variables) aggregate(cbind(Murder, Assault, UrbanPop, Rape) ~ ID, data = df, FUN = mean, na.rm = TRUE) # Or sum a specific column, like total Assaults per ID aggregate(Assault ~ ID, data = df, FUN = sum)
Group by Multiple Variables (ID + Year)
Count observations and compute median Rape rates for each ID-Year combination:
# Number of rows per ID-Year group aggregate(Murder ~ ID + Year, data = df, FUN = length) # Using Murder as a placeholder column # Median Rape rate per ID-Year group aggregate(Rape ~ ID + Year, data = df, FUN = median, na.rm = TRUE)
2. dplyr/Tidyverse Approach
This is the more flexible and readable method for complex grouping tasks. First, load the package if you haven't already:
library(dplyr)
Group by ID and Compute Custom Aggregates
Use group_by() to define groups, then summarize() to calculate stats:
df %>% group_by(ID) %>% summarize( avg_murder = mean(Murder, na.rm = TRUE), total_assault = sum(Assault, na.rm = TRUE), num_records = n(), # Count rows in each group max_urbanpop = max(UrbanPop, na.rm = TRUE) ) %>% ungroup() # Always ungroup after operations to avoid unexpected behavior later
Group by ID + Year
Easily compute multiple stats across groups with across():
df %>% group_by(ID, Year) %>% summarize( across(c(Murder, Rape), mean, na.rm = TRUE), # Mean of Murder/Rape per group total_urbanpop = sum(UrbanPop, na.rm = TRUE) ) %>% ungroup()
Grouped Transformation (Add Stats to Original Data)
If you want to append group-level stats to every row (e.g., average Murder rate for each ID):
df_with_group_stats <- df %>% group_by(ID) %>% mutate(avg_murder_for_id = mean(Murder, na.rm = TRUE)) %>% ungroup() # View the result head(df_with_group_stats)
Bonus: Useful Aggregation Functions
Here are some common functions you can use in summarize() or aggregate():
n(): Count rows in a groupsum(),mean(),median(): Summarize numeric valuesmax(),min(): Get extreme valuessd(): Calculate standard deviationfirst(),last(): Retrieve the first/last value in a group
内容的提问来源于stack exchange,提问作者WoeIs

