如何用R语言按区域及自定义年龄分组汇总客户出现次数?
问题:生成包含零计数的分组汇总表
现有包含Area和CustomerAge两列的数据集,已通过以下tidyverse代码生成客户年龄分组Customer_Age_Group:
library(tidyverse) df <- read.table(textConnection("Area CustomerAge A 28 A 40 A 70 A 19 B 13 B 12 B 72 B 90"), header=TRUE) df2 <- df %>% mutate( # 创建年龄分组 Customer_Age_Group = dplyr::case_when( CustomerAge <= 18 ~ "0-18", CustomerAge > 18 & CustomerAge <= 60 ~ "19-60", CustomerAge > 60 ~ ">60" ))
需求是按区域(Area)、**客户年龄分组(Customer_Age_Group)**汇总出现次数(Occurrences),且必须包含出现次数为0的分组,最终目标格式如下:
| 区域(Area) | 客户年龄分组(Customer_Age_Group) | 出现次数(Occurrences) |
|---|---|---|
| A | 0-18 | 0 |
| A | 19-60 | 3 |
| A | >60 | 1 |
| B | 0-18 | 2 |
| B | 19-60 | 0 |
| B | >60 | 2 |
解决方案
要补全零计数的分组,核心是先构建Area与所有年龄分组的完整组合,再和原始分组数据做关联统计:
# 定义所有需要覆盖的年龄分组 age_groups <- c("0-18", "19-60", ">60") # 生成区域与年龄分组的全量组合 full_combinations <- expand_grid( Area = unique(df$Area), Customer_Age_Group = age_groups ) # 汇总计数并补全零值 result <- df2 %>% count(Area, Customer_Age_Group, name = "Occurrences") %>% right_join(full_combinations, by = c("Area", "Customer_Age_Group")) %>% mutate(Occurrences = replace_na(Occurrences, 0)) %>% arrange(Area, Customer_Age_Group) # 输出结果 result
运行上述代码后,即可得到符合要求的完整汇总表。
内容的提问来源于stack exchange,提问作者JeffWithpetersen
相关产品推荐
相关产品推荐

