R语言如何按年份及指定阈值条件统计dataframe数据行数
R语言按年份统计DataFrame数据点实现方案
先确认Date列已转为标准日期类型,避免年份提取出错,若未转换先执行:df$Date <- as.Date(df$Date)
你之前写的年份列生成代码df$Year <- format(df$Date, format="%Y")逻辑正确,后续统计直接基于生成的Year列开展即可。
1. 统计各年份总数据点数
基础R实现
# 生成各年份总计数表 year_total <- as.data.frame(table(df$Year)) colnames(year_total) <- c("Year", "total_count")
dplyr实现(推荐大型数据集使用,运行效率更高)
library(dplyr) year_total <- df %>% group_by(Year) %>% summarise(total_count = n(), .groups = "drop")
2. 统计各年份Weight高于指定阈值的点数
先把需要统计的阈值统一存入向量,避免重复编写逻辑:threshold_list <- c(0.1, 0.2, 0.5)
基础R实现
# 以总计数表为基础合并各阈值统计结果 stat_result <- year_total for (th in threshold_list) { # 筛选当前阈值以上的子集,按年份计数 sub_count <- aggregate( Weight ~ Year, data = df[df$Weight > th, ], FUN = length ) colnames(sub_count)[2] <- paste0("count_gt_", th) stat_result <- merge(stat_result, sub_count, by = "Year", all.x = TRUE) } # 无符合条件数据的年份对应值替换为0 stat_result[is.na(stat_result)] <- 0
dplyr实现
library(dplyr) stat_result <- df %>% group_by(Year) %>% summarise( total_count = n(), # 批量生成所有阈值的统计列 across( .cols = all_of(threshold_list), .fns = ~ sum(Weight > .x), .names = "count_gt_{.col}" ), .groups = "drop" )
样例输出
基于你给出的示例数据,最终统计结果如下:
| Year | total_count | count_gt_0.1 | count_gt_0.2 | count_gt_0.5 |
|---|---|---|---|---|
| 2018 | 2 | 1 | 1 | 0 |
| 2019 | 2 | 2 | 2 | 1 |
| 2020 | 2 | 2 | 2 | 0 |
后续绘制柱状图时,直接调用stat_result表中对应字段映射绘图参数即可。
内容的提问来源于stack exchange,提问作者alec22
相关产品推荐
相关产品推荐

