You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为什么dplyr::tally()会移除输出中的最后一个分组变量?

关于dplyr::tally()移除最后一个分组变量的说明

这确实是dplyr的预期设计行为,原因和tally()的定位直接相关:

  • tally()是summarise()的便捷封装函数,核心作用是对当前最细粒度的分组做计数汇总。当你对多层分组的数据框调用它时,它会针对最后一层(最细分的)分组维度完成计数,这层分组在汇总后就没有继续保留分组标记的必要了,所以会被自动移除。
  • 这种设计贴合常规数据处理逻辑:完成最细分组的计数后,后续操作通常是基于上层分组展开,保留前面的分组维度更符合实际使用场景。

关于这个行为的官方说明,可以在R中执行?dplyr::tally查看文档,其中明确提到:

tally() is a convenient wrapper for summarise that will either call n() or sum(n) depending on whether you're tallying for the first time, and will ungroup the last group by default (since you've usually counted across it).

你提供的代码示例也验证了这个行为:

library(dplyr) # 1.0.10

# single grouping variable returns ungrouped output
starwars %>%
  group_by(species) %>%
  tally() %>%
  groups()

#> list()

# three grouping variables return output grouped by first and second group
starwars %>%
  group_by(species, eye_color, skin_color) %>%
  tally() %>%
  groups()

#> [[1]]
#> species
#>
#> [[2]]
#> eye_color

内容的提问来源于stack exchange,提问作者Patrick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 01:35:56