R语言中如何按loc.id和year分组计算DataFrame列均值?
Hey there! Let's break down how to calculate column means grouped by loc.id and year for your dataset. I'll cover three popular approaches so you can pick what fits your workflow best:
1. Base R Approach (No External Packages)
If you don't want to install extra packages, base R's aggregate() function gets the job done cleanly. It lets you group by multiple variables and apply a function to all other columns:
# First recreate your dataset (as provided) set.seed(123) df <- data.frame(loc.id = rep(1:5, each = 4*2), year = rep(rep(1980:1983,each = 2),times = 5), type = rep(2:3, times = 4*5), x = runif(5*4*2), y = runif(5*4*2), z = runif(5*4*2)) # Calculate grouped means with base R grouped_means_base <- aggregate(. ~ loc.id + year, data = df, FUN = mean) # Verify the subset you mentioned (loc.id=1, year=1980) subset(grouped_means_base, loc.id == 1 & year == 1980)
- The
. ~ loc.id + yearsyntax tells R to group byloc.idandyear, then apply themeanfunction to all other columns (includingtype,x,y,z). - The
subset()line lets you quickly check the specific group you asked about to confirm results.
2. Tidyverse (dplyr) Approach
If you prefer a more readable, pipe-based workflow, the dplyr package (part of the tidyverse) is perfect. It uses explicit verbs to make your code easy to follow:
# Load dplyr first (run install.packages("dplyr") if you haven't installed it) library(dplyr) grouped_means_dplyr <- df %>% group_by(loc.id, year) %>% # Define your grouping variables summarise( across(c(x, y, z), mean), # Calculate mean for x, y, z specifically .groups = "drop" # Optional: removes grouping structure from the final output ) # Check the target subset grouped_means_dplyr %>% filter(loc.id == 1 & year == 1980)
group_by()sets up your groups,summarise()withacross()lets you specify exactly which columns to compute means for (useeverything()instead ofc(x,y,z)if you want all non-group columns included).- The
.groups = "drop"argument ensures you get a regular data frame back instead of a grouped one (skip this if you want to keep grouping for further operations).
3. Data.table Approach (Fast for Large Datasets)
If you're working with big data, data.table is the go-to for speed and concise syntax. It's especially efficient for large grouped operations:
# Load data.table first (run install.packages("data.table") if needed) library(data.table) # Convert your data frame to a data.table in-place setDT(df) # Calculate grouped means grouped_means_dt <- df[, lapply(.SD, mean), by = .(loc.id, year)] # Check the target subset grouped_means_dt[loc.id == 1 & year == 1980]
setDT()converts your existing data frame to a data.table without making a copy.by = .(loc.id, year)defines the groups, andlapply(.SD, mean)applies the mean function to every column in.SD(the "subset of data" excluding the grouping columns).
All three methods will give you the exact same mean values for each loc.id + year group—pick the one that aligns with your current R workflow!
内容的提问来源于stack exchange,提问作者89_Simple

