R dplyr循环多变量分组summarise合并结果出错怎么解决
问题根源
你编写的循环代码错误出在group_by的非标准求值逻辑上:groups[i]是字符串类型,直接用!!展开相当于按字符串常量分组,整个数据集只会生成1个取值为变量名的分组,最终统计的是全量数据集的汇总值,所以才会出现3行全量汇总的错误结果。要让dplyr识别到你需要按字符串对应的数据集变量分组,需要先把字符串转为R的符号对象。
修正后的循环代码
只需要把group_by(!!groups[i])改为group_by(!!sym(groups[i]))即可:
groups <- c("sexe", "trav.satisf", "cuisine") synthese <- tibble() for (i in seq_along(groups)) { tmp <- hdv2003 %>% group_by(!!sym(groups[i])) %>% summarise(n = n(), percent = round((n()/nrow(hdv2003))*100, digits = 1), femmes = round((sum(sexe == "Femme", na.rm = TRUE)/sum(!is.na(sexe)))*100, digits = 1), age = round(mean(age, na.rm = TRUE), digits = 1) ) names(tmp)[1] <- "group" synthese <- bind_rows(synthese, tmp) }
运行上述代码即可得到和你手动多次分组拼接完全一致的结果。
更简洁的purrr实现(可选)
你也可以用purrr包的迭代函数替代for循环,代码更简洁,不需要手动维护结果对象:
library(purrr) groups <- c("sexe", "trav.satisf", "cuisine") synthese <- map_dfr(groups, function(var) { hdv2003 %>% group_by(!!sym(var)) %>% summarise( n = n(), percent = round((n()/nrow(hdv2003))*100, 1), femmes = round((sum(sexe == "Femme", na.rm = T)/sum(!is.na(sexe)))*100, 1), age = round(mean(age, na.rm = T), 1) ) %>% rename(group = 1) })
内容的提问来源于stack exchange,提问作者Joël
相关产品推荐
相关产品推荐

