You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R dplyr循环多变量分组summarise合并结果出错怎么解决

问题根源

你编写的循环代码错误出在group_by的非标准求值逻辑上:groups[i]是字符串类型,直接用!!展开相当于按字符串常量分组,整个数据集只会生成1个取值为变量名的分组,最终统计的是全量数据集的汇总值,所以才会出现3行全量汇总的错误结果。要让dplyr识别到你需要按字符串对应的数据集变量分组,需要先把字符串转为R的符号对象。

修正后的循环代码

只需要把group_by(!!groups[i])改为group_by(!!sym(groups[i]))即可:

groups <- c("sexe",
            "trav.satisf",
            "cuisine")

synthese <- tibble()

for (i in seq_along(groups)) {
  tmp <- hdv2003 %>%
    group_by(!!sym(groups[i])) %>%  
    summarise(n = n(),
              percent = round((n()/nrow(hdv2003))*100, digits = 1),
              femmes = round((sum(sexe == "Femme", na.rm = TRUE)/sum(!is.na(sexe)))*100, digits = 1),
              age = round(mean(age, na.rm = TRUE), digits = 1)
    )
  
  names(tmp)[1] <- "group"
  synthese <- bind_rows(synthese, tmp)
}

运行上述代码即可得到和你手动多次分组拼接完全一致的结果。

更简洁的purrr实现(可选)

你也可以用purrr包的迭代函数替代for循环,代码更简洁,不需要手动维护结果对象:

library(purrr)
groups <- c("sexe", "trav.satisf", "cuisine")

synthese <- map_dfr(groups, function(var) {
  hdv2003 %>%
    group_by(!!sym(var)) %>%
    summarise(
      n = n(),
      percent = round((n()/nrow(hdv2003))*100, 1),
      femmes = round((sum(sexe == "Femme", na.rm = T)/sum(!is.na(sexe)))*100, 1),
      age = round(mean(age, na.rm = T), 1)
    ) %>%
    rename(group = 1)
})

内容的提问来源于stack exchange,提问作者Joël

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 23:54:03