R中合并频数表分组计数的报错排查及通用实现方案问询
问题描述
我在R中创建了如下年度样本频数表:
1975 1976 1980 1995 2017 2018 2 67 44 126 64 133
需要将部分年份的计数合并,目标结果如下:
Eighties Nineties Two Thousands 2+67+44 126 64+133
尝试过的方法及报错
嵌套ifelse函数
模仿Excel逻辑写的嵌套ifelse代码:
a$decade -> ifelse(a$year=1975 | a$year=1976 |a$year=1980,"Eighties", ifelse(a$year=1995,"Ninties", ifelse(a$year=2017|a$year=2018,"Two Thousands","Error")))
运行后报错:
Error: unexpected '=' in "a$decade -> ifelse(a$year="E > ifelse(a$year=1995,"Ninties", Error: unexpected '=' in " ifelse(a$year=" > ifelse(a$year=2017|a$year=2018,"Two Thousands","Error"))) Error: unexpected '=' in " ifelse(a$year=" >
错误的多条件ifelse写法
随后尝试的代码同样报错:
a$decade <- ifelse(a$year %in% (c("1975", "1976","1980"), "Eighties", (c("1995"),"Nineties"),(c("2017", "2018"),"2000s"))
dplyr的case_when分组求和
最后尝试用dplyr分组求和:
a %>% mutate(decade = case_when( year %in% c(1975, 1976, 1980) ~ "Eighties", year == 1995 ~ "Nineties", year %in% c(2017, 2018) ~ "Two Thousands" )) %>% group_by(decade) %>% summarize(count = sum(count)) %>% ungroup()
报错信息:
Error: Problem with `summarise()` column `count`. i `count = sum(count)`. x invalid 'type' (closure) of argument i The error occurred in group 1: decade = "Eighties".
注:该逻辑需复用至其他场景,如月份分季节、日期分工作日/周末、采样点分方位等,希望得到通用方案。
报错原因及解决方案
各报错原因解析
- 嵌套ifelse报错:R中判断相等必须用
==,而非=(=在函数参数中是赋值作用);同时存在拼写错误Ninties,应为Nineties。 - 多条件ifelse报错:
ifelse()仅接受3个参数(条件、满足值、不满足值),无法直接传入多组(条件,值)对,语法逻辑错误。 - dplyr代码报错:数据框中没有名为
count的列,R将count识别为dplyr的count()函数(类型为closure),导致求和参数类型错误。
修正后的代码及通用方案
首先确保数据框结构正确,示例数据:
a <- data.frame( year = c(1975, 1976, 1980, 1995, 2017, 2018), freq = c(2, 67, 44, 126, 64, 133) )
方案1:修正嵌套ifelse(适合简单规则)
a$decade <- ifelse(a$year == 1975 | a$year == 1976 | a$year == 1980, "Eighties", ifelse(a$year == 1995, "Nineties", ifelse(a$year == 2017 | a$year == 2018, "Two Thousands", "Error"))) # 分组求和 aggregate(freq ~ decade, data = a, sum)
方案2:dplyr case_when(适合直观的多条件规则)
library(dplyr) a %>% mutate(decade = case_when( year %in% c(1975, 1976, 1980) ~ "Eighties", year == 1995 ~ "Nineties", year %in% c(2017, 2018) ~ "Two Thousands", TRUE ~ "Unclassified" # 兜底处理未匹配项 )) %>% group_by(decade) %>% summarise(total = sum(freq)) %>% ungroup()
方案3:映射表匹配(最通用,适合复用至其他场景)
通过定义分组映射表,可以快速将规则复用至月份分季节、采样点分方位等场景:
# 定义年度分组映射表 group_mapping <- data.frame( year = c(1975,1976,1980,1995,2017,2018), decade = c(rep("Eighties",3), "Nineties", rep("Two Thousands",2)) ) # 示例:月份分季节的映射表(直接替换即可复用) # season_mapping <- data.frame( # month = 1:12, # season = c(rep("Winter",2), rep("Spring",3), rep("Summer",3), rep("Autumn",3), "Winter") # ) # 合并分组并求和 a %>% left_join(group_mapping, by = "year") %>% group_by(decade) %>% summarise(total = sum(freq)) %>% ungroup()
方案4:factor重编码(适合固定值的分组)
a$decade <- factor(a$year, levels = c(1975,1976,1980,1995,2017,2018), labels = c(rep("Eighties",3), "Nineties", rep("Two Thousands",2))) a %>% group_by(decade) %>% summarise(total = sum(freq)) %>% ungroup()
内容的提问来源于stack exchange,提问作者creusac
相关产品推荐
相关产品推荐

