You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何summarise中mean()可用count()报错?如何按月统计绘制柱状图?

问题:dplyr中summarise里mean可用但count/n报错,如何正确绘制每月行数统计的柱状图

问题背景

当前用于生成柱状图的代码统计结果不正确:

# counts per month
df %>%
  group_by(date) %>%
  summarize(dep_delay= mean(dep_delay)) %>%
  ggplot(aes(x = as.factor(month(date,label = TRUE)),
             y = dep_delay)) +
  geom_bar(stat = 'identity') +
  theme_bw()+
  scale_y_continuous(breaks = seq(0,1400,200))+
  labs(title="Departure delays in Newark airport per month in 2013", 
       x="", y = "number of delays")

正确的按月统计行数结果应为:

df_w_delays %>% group_by(month=month(date)) %>% count()

# 输出结果
  month     n
   <dbl> <int>
 1     1  1035
 2     2   941
 3     3  1239
 4     4  1159
 5     5  1240
 6     6  1257
 7     7  1277
 8     8  1100
 9     9   702
10    10   880
11    11   861
12    12  1315

修改summarize统计逻辑时出现以下报错:

  1. 使用n(dep_delay)时:
Error in `summarize()`:
! Problem while computing `dep_delay = n(dep_delay)`.
ℹ The error occurred in group 1: date = 2013-01-01.
Caused by error in `n()`:
! unused argument (dep_delay)
  1. 使用count(dep_delay)时:
Error in `summarize()`:
! Problem while computing `dep_delay = count(dep_delay)`.
ℹ The error occurred in group 1: date = 2013-01-01.
Caused by error in `UseMethod()`:
! no applicable method for 'count' applied to an object of class "c('double', 'numeric')"

为什么mean()可以在summarise中正常使用,而count()/n()不行?

  • mean()是向量函数:它接受一列向量(比如dep_delay)作为参数,计算该列的均值,完全符合summarise对列做聚合运算的逻辑。
  • n()是dplyr专用聚合函数:它不需要传入参数,作用是返回当前分组的行数,不能给它加列名参数,所以n(dep_delay)会报错。
  • count()是dplyr顶层函数:它是直接作用于整个数据框的工具(比如df %>% count(col)),不是用于summarise内部的聚合函数,不能在summarise里直接传入单个列调用,因此count(dep_delay)会报错。

如何在柱状图中展示每月的行数统计?

核心是先按月份正确分组统计行数,再绘图,以下是两种可行方式:

方式1:用group_by + summarise(n())

df %>%
  # 提取月份并按月份分组
  group_by(month = month(date, label = TRUE)) %>%
  # 统计每组行数,命名为dep_delay适配原绘图的y轴
  summarize(dep_delay = n()) %>%
  ggplot(aes(x = month, y = dep_delay)) +
  geom_bar(stat = 'identity') +
  theme_bw()+
  scale_y_continuous(breaks = seq(0,1400,200))+
  labs(title="Departure delays in Newark airport per month in 2013", 
       x="", y = "number of delays")

方式2:用count()直接生成统计结果

df %>%
  # 按月份分组统计行数,指定列名为dep_delay
  count(month = month(date, label = TRUE), name = "dep_delay") %>%
  ggplot(aes(x = month, y = dep_delay)) +
  geom_bar(stat = 'identity') +
  theme_bw()+
  scale_y_continuous(breaks = seq(0,1400,200))+
  labs(title="Departure delays in Newark airport per month in 2013", 
       x="", y = "number of delays")

原代码的核心问题

你之前是按date(具体日期)分组,之后再把日期转成月份聚合,这会先计算每日均值,再把每日均值当作月度数值,完全偏离了统计月度行数的需求。


内容的提问来源于stack exchange,提问作者Bluetail

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 13:55:12