dplyr的group_by函数无法正常分组,聚合后仅返回全局均值如何解决
问题产生原因
dplyr的分组属性是和数据框中的列绑定的,你在group_by(species, sex)之后调用select()仅保留了height和mass两列,直接移除了作为分组依据的species和sex列,分组属性也会随之丢失。后续summarise()执行时没有识别到分组规则,就会对全量数据做聚合,最终仅返回一行全局均值结果。
解决方法
有两种常用的修复方案,任选其一即可:
- 方案1:调整函数执行顺序,先做列筛选再执行分组,避免分组列被误删
starwars %>% select(species, sex, height, mass) %>% group_by(species, sex) %>% summarise( height = mean(height, na.rm = TRUE), mass = mean(mass, na.rm = TRUE), .groups = "drop" # 可选,聚合后自动取消分组,避免残留分组影响后续操作 )
- 方案2:
select()时保留分组列,不需要调整原有执行顺序
starwars %>% group_by(species, sex) %>% select(species, sex, height, mass) %>% summarise( height = mean(height, na.rm = TRUE), mass = mean(mass, na.rm = TRUE) )
内容的提问来源于stack exchange,提问作者int12345
相关产品推荐
相关产品推荐

