You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按多分组变量同时统计两个因子列的出现次数?

解决方案

首先准备可复现的示例数据集用于测试:

set.seed(123)
library(dplyr)

df <- tibble(
  year = rep(2020:2022, each = 10),
  temp_catog = sample(c("low", "mid", "high"), 30, replace = TRUE),
  humidity_catog = sample(c("dry", "normal", "wet"), 30, replace = TRUE)
)

如果要基于year + humidity_catog + temp_catog的多分组逻辑,同时得到每一年中humidity_catog各分类的总出现次数(rh_count)、temp_catog各分类的总出现次数(temp_count),可以用add_count函数快速实现,代码如下:

result <- df %>%
  # 按年份+湿度分类统计总次数,命名为rh_count
  add_count(year, humidity_catog, name = "rh_count") %>%
  # 按年份+温度分类统计总次数,命名为temp_count
  add_count(year, temp_catog, name = "temp_count") %>%
  # 保留唯一的(year, humidity_catog, temp_catog)组合及对应的计数
  distinct(year, humidity_catog, temp_catog, rh_count, temp_count) %>%
  ungroup()

代码说明

  • add_count()会自动为指定分组添加计数列,无需手动写group_by()+summarize(),比单独分组统计再合并更高效。
  • distinct()用来去除重复的组合行,确保每一行都是唯一的年份+湿度分类+温度分类组合,同时保留对应的两类计数。
  • 最终结果里,rh_count是该湿度分类在对应年份的总出现次数,temp_count是该温度分类在对应年份的总出现次数,完全符合“统计实际出现次数”的要求。

如果一定要显式使用group_by(year, humidity_catog, temp_catog)的多分组起始逻辑,也可以用嵌套分组的方式实现:

result <- df %>%
  group_by(year, humidity_catog, temp_catog) %>%
  # 标记当前组合的出现次数(可选)
  mutate(combination_count = n()) %>%
  # 切换分组:仅按年份+湿度分类统计总次数
  group_by(year, humidity_catog, .add = FALSE) %>%
  mutate(rh_count = n()) %>%
  # 再次切换分组:按年份+温度分类统计总次数
  group_by(year, temp_catog, .add = FALSE) %>%
  mutate(temp_count = n()) %>%
  # 去重得到唯一组合结果
  distinct(year, humidity_catog, temp_catog, rh_count, temp_count, combination_count) %>%
  ungroup()

这个写法完全遵循了group_by(year, humidity_catog, temp_catog)的多分组起始逻辑,通过.add = FALSE切换分组维度来计算不同分类的年度总次数。

内容的提问来源于stack exchange,提问作者Ahsk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 17:23:26