如何在R语言中按年份聚合数据生成DataFrame?
按年份聚合列表数据生成指定结构的DataFrame
你需要将包含2011、2012年月度数据的列表,按年份聚合为以时间粒度(如m5、m10)为行名、年份为列的DataFrame,这里提供两种简洁的实现方式:
样本数据确认
首先先明确你的输入数据结构(和你提供的一致):
data <- list( `01-2011` = structure(c(0.266, 0.532, 0.797, 1.092, 1.27, 1.27, 1.27, 1.46, 1.46, 2.34, 2.53, 2.53, 2.53, 2.53), .Dim = c(14L, 1L), .Dimnames = list(c("m5", "m10", "m15", "m30", "h1", "h2", "h3", "h4", "h5", "h6", "h8", "h12", "h18", "h24"), NULL)), `02-2011` = structure(c(0.955, 1.683, 2.398, 4.539, 6.528, 9.427, 10.848, 9.543, 13.736, 16.635, 16.751, 16.751, 16.751, 16.751), .Dim = c(14L, 1L), .Dimnames = list(c("m5", "m10", "m15", "m30", "h1", "h2", "h3", "h4", "h5", "h6", "h8", "h12", "h18", "h24"), NULL)), `01-2012` = structure(c(1.224, 2.395, 3.063, 5.131, 7.112, 9.474, 9.474, 10.302, 10.744, 9.474, 12.49, 11.406, 13.571, 13.919), .Dim = c(14L, 1L), .Dimnames = list(c("m5", "m10", "m15", "m30", "h1", "h2", "h3", "h4", "h5", "h6", "h8", "h12", "h18", "h24"), NULL)), `03-2012` = structure(c(0.75, 1.391, 1.871, 3.649, 5.174, 6.275, 6.439, 8.396, 6.963, 10.453, 8.844, 10.453, 10.901, 10.901), .Dim = c(14L, 1L), .Dimnames = list(c("m5", "m10", "m15", "m30", "h1", "h2", "h3", "h4", "h5", "h6", "h8", "h12", "h18", "h24"), NULL)) )
方案1:基础R原生实现(无需额外包)
用基础R的函数就能完成,步骤清晰:
# 1. 按年份对列表中的元素分组 year_groups <- split(data, substr(names(data), 4, 7)) # 2. 对每个年份下的所有月度数据,按行执行聚合操作(这里用sum,你可以换成mean等) aggregated_data <- lapply(year_groups, function(month_data) { rowSums(do.call(cbind, month_data)) }) # 3. 转换为DataFrame并设置行名为时间粒度 output <- as.data.frame(aggregated_data) row.names(output) <- rownames(data[[1]])
方案2:tidyverse风格实现(更直观)
如果你习惯使用dplyr、purrr这类工具,这种写法更符合现代R的编程风格:
library(dplyr) library(purrr) library(tibble) library(tidyr) output <- data %>% # 将列表转换为长格式数据框,保留时间粒度、数值和年月标识 imap_dfr(~tibble(time = rownames(.x), value = .x[,1], year_month = .y)) %>% # 从年月中提取年份 mutate(year = substr(year_month, 4, 7)) %>% # 按时间粒度和年份分组聚合(这里用sum,可按需替换) group_by(time, year) %>% summarise(total = sum(value), .groups = "drop") %>% # 转换为宽格式,年份作为列 pivot_wider(names_from = year, values_from = total) %>% # 将时间粒度列设置为行名 column_to_rownames(var = "time")
查看输出结果
运行上述任意一种方案后,你就能得到符合需求的DataFrame,比如查看前6行的结果:
head(output)
输出示例:
2011 2012 m5 1.221 1.974 m10 2.215 3.786 m15 3.195 4.934 m30 5.631 8.780 h1 7.798 12.286 h2 10.697 15.749
内容的提问来源于stack exchange,提问作者Hüsamettin Tayşi
相关产品推荐
相关产品推荐

