R语言:按site分组聚合N列求和并将结果重复填充至新列的实现方案
Hey there! I totally get figuring out grouped calculations in R when you're just starting out—let's walk through a couple of easy ways to get your desired result.
dplyr (Tidyverse approach) This is super intuitive if you're getting into the tidyverse ecosystem. First, make sure you have the package installed (if not, run install.packages("dplyr")), then use group_by() to split your data by site, and mutate() to add the new column with the group sum:
# Load the package library(dplyr) # Your original data site = c('a', 'a', 'a', 'b', 'b', 'b', 'b', 'b', 'c', 'c', 'c', 'c') N = c(5, 4, 2, 5, 10, 15, 6, 4, 29, 14, 10, 12) df = data.frame(site, N) # Add the grouped sum column df <- df %>% group_by(site) %>% mutate(N_total_site = sum(N)) %>% ungroup() # Optional, but good practice to remove grouping # Check the result print(df)
When you run this, mutate() keeps all the original rows and fills each group's sum into every row of that group—exactly what you want!
If you don't want to load any extra packages, the ave() function is perfect for this scenario. It calculates grouped aggregates and returns a vector with the same length as your original data, repeating the group result for each row in the group:
# Your original data site = c('a', 'a', 'a', 'b', 'b', 'b', 'b', 'b', 'c', 'c', 'c', 'c') N = c(5, 4, 2, 5, 10, 15, 6, 4, 29, 14, 10, 12) df = data.frame(site, N) # Add the grouped sum column df$N_total_site <- ave(df$N, df$site, FUN = sum) # Check the result print(df)
This is a concise, base-R-only solution that works great for small to medium datasets.
data.table (For large datasets) If you're working with big data later on, data.table is lightning-fast. Here's how to do it with this package:
# Install and load the package if needed # install.packages("data.table") library(data.table) # Convert your data frame to a data.table dt <- as.data.table(df) # Add the grouped sum column (in-place, which is efficient) dt[, N_total_site := sum(N), by = site] # Convert back to data frame if needed df <- as.data.frame(dt) # Check the result print(df)
This method modifies the data in place (using :=), which saves memory—great for large datasets.
All three methods will give you the exact result you're looking for:
site N N_total_site
1 a 5 11
2 a 4 11
3 a 2 11
4 b 5 40
5 b 10 40
6 b 15 40
7 b 6 40
8 b 4 40
9 c 29 65
10 c 14 65
11 c 10 65
12 c 12 65
内容的提问来源于stack exchange,提问作者hiperhiper

