You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言如何按年份及指定阈值条件统计dataframe数据行数

R语言按年份统计DataFrame数据点实现方案

先确认Date列已转为标准日期类型,避免年份提取出错,若未转换先执行:
df$Date <- as.Date(df$Date)
你之前写的年份列生成代码df$Year <- format(df$Date, format="%Y")逻辑正确,后续统计直接基于生成的Year列开展即可。


1. 统计各年份总数据点数

基础R实现

# 生成各年份总计数表
year_total <- as.data.frame(table(df$Year))
colnames(year_total) <- c("Year", "total_count")

dplyr实现(推荐大型数据集使用,运行效率更高)

library(dplyr)
year_total <- df %>%
  group_by(Year) %>%
  summarise(total_count = n(), .groups = "drop")

2. 统计各年份Weight高于指定阈值的点数

先把需要统计的阈值统一存入向量,避免重复编写逻辑:
threshold_list <- c(0.1, 0.2, 0.5)

基础R实现

# 以总计数表为基础合并各阈值统计结果
stat_result <- year_total
for (th in threshold_list) {
  # 筛选当前阈值以上的子集,按年份计数
  sub_count <- aggregate(
    Weight ~ Year, 
    data = df[df$Weight > th, ], 
    FUN = length
  )
  colnames(sub_count)[2] <- paste0("count_gt_", th)
  stat_result <- merge(stat_result, sub_count, by = "Year", all.x = TRUE)
}
# 无符合条件数据的年份对应值替换为0
stat_result[is.na(stat_result)] <- 0

dplyr实现

library(dplyr)
stat_result <- df %>%
  group_by(Year) %>%
  summarise(
    total_count = n(),
    # 批量生成所有阈值的统计列
    across(
      .cols = all_of(threshold_list),
      .fns = ~ sum(Weight > .x),
      .names = "count_gt_{.col}"
    ),
    .groups = "drop"
  )

样例输出

基于你给出的示例数据,最终统计结果如下:

Yeartotal_countcount_gt_0.1count_gt_0.2count_gt_0.5
20182110
20192221
20202220

后续绘制柱状图时,直接调用stat_result表中对应字段映射绘图参数即可。

内容的提问来源于stack exchange,提问作者alec22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 00:03:20