如何在R中基于年积日(DOY)对多列数据进行每日求和?
Hey Sarah, I see where you hit a snag with the aggregate() function—let’s get this sorted out for you! The issue was your column selection syntax, and we have a couple of straightforward fixes to sum columns 4 through 44 by DOY (or Year+DOY, if you need to account for overlapping DOYs across different years).
Method 1: Using Formula Syntax with Column Subsetting
First, we’ll explicitly select the grouping variable(s) and the columns you want to sum, then use . in the formula to represent all non-grouping columns. We’ll also add na.rm = TRUE to handle those NaN values in your data (otherwise, your sums would return NaN).
Code for DOY-only grouping:
# Subset to DOY + columns 4-44, then aggregate Js_gm2d <- aggregate(. ~ DOY, data = Js_30min[, c("DOY", 4:44)], FUN = sum, na.rm = TRUE)
Code for Year+DOY grouping (to avoid merging same DOYs across different years):
# Subset to Year + DOY + columns 4-44, then aggregate Js_gm2d <- aggregate(. ~ Year + DOY, data = Js_30min[, c("Year", "DOY", 4:44)], FUN = sum, na.rm = TRUE)
Method 2: Using x and by Parameters (More Explicit)
If you prefer a more direct approach without formula syntax, you can specify the columns to aggregate (x) and the grouping variables (by) separately:
Code for DOY-only grouping:
Js_gm2d <- aggregate(x = Js_30min[, 4:44], by = list(DOY = Js_30min$DOY), FUN = sum, na.rm = TRUE)
Code for Year+DOY grouping:
Js_gm2d <- aggregate(x = Js_30min[, 4:44], by = list(Year = Js_30min$Year, DOY = Js_30min$DOY), FUN = sum, na.rm = TRUE)
Why Your Original Code Failed
Your line aggregate(c(,4:44)~DOY,data=Js_30min,FUN=sum) had invalid syntax: c(,4:44) doesn’t make sense to R—you need to explicitly define the columns you want to include, either by position, name, or by subsetting the data first.
Quick Test with Your Sample Data
Let’s verify with your sample data snippet:
# Sample data Js_30min <- data.frame( Year = rep(2018, 6), DOY = rep(157, 6), Time = c(0, 30, 100, 130, 200, 230), SI1 = c(0, 297.9779, 349.7710, 101.0535, 143.9961, 0), SI2 = c(0, 493.8855, 1168.0555, 1279.5865, 1392.9739, 891.1722), SI3 = rep(NaN, 6) ) # Run aggregation Js_gm2d <- aggregate(. ~ DOY, data = Js_30min[, c("DOY", 4:6)], FUN = sum, na.rm = TRUE)
This will return:
DOY SI1 SI2 SI3 1 157 892.7985 5225.674 0
Note that SI3 sums to 0 because all values were NaN and we used na.rm = TRUE.
内容的提问来源于stack exchange,提问作者Sarah

