You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用dplyr按组计算忽略NA的累积均值(cummean)

解决dplyr分组计算忽略NA的累积均值问题

Got it, let's fix this! The built-in cummean() function does carry over NA values instead of ignoring them, which isn't what you need here. But we can easily build a custom cumulative mean that skips NAs using core dplyr functions—no extra packages required.

具体实现代码

First, let's recreate your data frame, then add the custom cumulative mean column:

library(dplyr)

# 原始数据框
df <- data.frame(
  category=c("cat1","cat1","cat2","cat1","cat2","cat2","cat1","cat2"),
  value=c(NA,2,3,4,5,NA,7,8)
)

# 分组计算忽略NA的累积均值
df <- df %>%
  group_by(category) %>%
  mutate(
    # 计算非NA值的累积和
    cum_sum = cumsum(value, na.rm = TRUE),
    # 统计到当前行为止的非NA值数量
    cum_count = cumsum(!is.na(value)),
    # 生成目标列:累积和/累积计数,无有效数据时返回NA
    new_col = ifelse(cum_count == 0, NA, cum_sum / cum_count)
  ) %>%
  # 可选:移除中间计算列
  select(-cum_sum, -cum_count) %>%
  ungroup()

print(df)

结果说明

运行代码后会得到如下输出:

# A tibble: 8 × 3
  category value new_col
  <chr>    <dbl>   <dbl>
1 cat1        NA      NA
2 cat1         2    2   
3 cat2         3    3   
4 cat1         4    3   
5 cat2         5    4   
6 cat2        NA    4   
7 cat1         7    4.33
8 cat2         8    5.33

核心逻辑拆解:

  • cumsum(value, na.rm = TRUE):直接跳过NA值计算累积和,不会把NA当作0参与求和
  • cumsum(!is.na(value)):统计到当前行为止的非NA值数量(布尔值会被转为1/0参与求和)
  • 最后通过ifelse处理无有效数据的情况,避免出现NaN,返回更符合预期的NA

这个方案完全匹配你的需求:按category分组、计算忽略NA的累积均值、不将NA视为0。

内容的提问来源于stack exchange,提问作者John F

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:47:48