You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言aggregate()函数2011年示例无法复现,返回NAN值求助

问题描述

复现2011年的示例脚本时,base R的aggregate()函数返回NA值而非预期的汇总结果,想知道是否需要使用aggregate()的更新版本或替代函数。

生成NA值的代码如下:

s1.no.present <- aggregate(s1s2.df$no.present[s1s2.df$sabap==-1], by=list(s1s2.df$month.n[s1s2.df$sabap==-1]),sum)[,2]
s1.no.cards <- aggregate(s1s2.df$no.cards[s1s2.df$sabap==-1], by=list(s1s2.df$month.n[s1s2.df$sabap==-1]),sum)[,2]
s2.no.present <- aggregate(s1s2.df$no.present[s1s2.df$sabap==1], by=list(s1s2.df$month.n[s1s2.df$sabap==1]),sum)[,2]
s2.no.cards <- aggregate(s1s2.df$no.cards[s1s2.df$sabap==1], by=list(s1s2.df$month.n[s1s2.df$sabap==1]),sum)[,2]

错误输出:

> tibble(s1.no.present)
# A tibble: 12 × 1
   s1.no.present
           <int>
 1            NA
 2            NA
 3            NA
 4            NA
 5            NA
 6            NA
 7            NA
 8            NA
 9            NA
10            NA
11            NA
12            NA
解决方案

1. 先排查数据基础问题

先确认筛选后的数据是否有效,避免因空数据集或全NA列导致求和返回NA:

# 检查sabap=-1的记录数量
nrow(s1s2.df[s1s2.df$sabap == -1,])
# 检查目标列是否全为NA
sum(is.na(s1s2.df$no.present[s1s2.df$sabap == -1]))
sum(is.na(s1s2.df$month.n[s1s2.df$sabap == -1]))

2. 优化aggregate()写法

原代码的子集索引写法易出错,建议先筛选数据框再聚合,同时添加na.rm=TRUE忽略NA值:

# 处理sabap=-1的情况
s1_sub <- subset(s1s2.df, sabap == -1)
s1_agg <- aggregate(cbind(no.present, no.cards) ~ month.n, data = s1_sub, sum, na.rm = TRUE)
s1.no.present <- s1_agg$no.present
s1.no.cards <- s1_agg$no.cards

# 处理sabap=1的情况
s2_sub <- subset(s1s2.df, sabap == 1)
s2_agg <- aggregate(cbind(no.present, no.cards) ~ month.n, data = s2_sub, sum, na.rm = TRUE)
s2.no.present <- s2_agg$no.present
s2.no.cards <- s2_agg$no.cards

3. 替代函数推荐

如果需要更灵活高效的方案,推荐以下工具:

  • dplyr(语法直观):
library(dplyr)

# 一次性计算所有汇总值
result <- s1s2.df %>%
  group_by(sabap, month.n) %>%
  summarise(
    no.present_sum = sum(no.present, na.rm = TRUE),
    no.cards_sum = sum(no.cards, na.rm = TRUE),
    .groups = "drop"
  )

# 提取所需变量
s1.no.present <- result %>% filter(sabap == -1) %>% pull(no.present_sum)
s1.no.cards <- result %>% filter(sabap == -1) %>% pull(no.cards_sum)
s2.no.present <- result %>% filter(sabap == 1) %>% pull(no.present_sum)
s2.no.cards <- result %>% filter(sabap == 1) %>% pull(no.cards_sum)
  • data.table(大数据场景更快):
library(data.table)
setDT(s1s2.df)

result <- s1s2.df[, .(
  no.present_sum = sum(no.present, na.rm = TRUE),
  no.cards_sum = sum(no.cards, na.rm = TRUE)
), by = .(sabap, month.n)]

# 提取变量
s1.no.present <- result[sabap == -1, no.present_sum]
s1.no.cards <- result[sabap == -1, no.cards_sum]
# 其余变量同理提取

内容的提问来源于stack exchange,提问作者Rion Lerm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 05:40:30