You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在dplyr的summarise中使用行值,实现类似data.table的效果?

解决dplyr分组计算时重复行与警告的正确方式

先看示例数据:

library(dplyr)
library(data.table)

df <- data.frame(
  id = c(1,1,2,2),
  x = 1:4
)

需求很明确:按id分组,计算每组x的均值后乘以该组的id值。

用data.table实现起来很直观:

setDT(df)[, .(mean(x) * id), by=id]

输出结果符合预期,每组一行:

id  V1
1:  1 1.5
2:  2 7.0

但用dplyr的summarise直接写时,会出现重复行和警告:

df |> group_by(id) |> summarise(mean(x) * id)

输出及警告:

`summarise()` has grouped output by 'id'. You can override using the `.groups` argument.
# A tibble: 4 × 2
# Groups:   id [2]
     id `mean(x) * id`
  <dbl>          <dbl>
1     1            1.5
2     1            1.5
3     2            7  
4     2            7  
Warning message:
Returning more (or less) than 1 row per `summarise()` group was deprecated in dplyr 1.1.0.
ℹ Please use `reframe()` instead.

虽然加unique()能去重,但这不是dplyr的正确用法,这里提供几种规范的解决方式:

方式一:用reframe()替代summarise()

dplyr 1.1.0之后推荐用reframe()处理非单行的分组结果,这里我们可以明确取组内的id值(因为组内id都相同,用first(id)即可):

df |> group_by(id) |> reframe(result = mean(x) * first(id))

输出:

# A tibble: 2 × 2
     id result
  <dbl>  <dbl>
1     1    1.5
2     2    7  

方式二:在summarise()中确保结果为标量

summarise()要求每组返回单行结果,所以需要把id转换成标量(组内id唯一,用first(id)或unique(id)都可以),同时可以加上.groups = "drop"取消分组状态:

df |> group_by(id) |> summarise(result = mean(x) * first(id), .groups = "drop")

输出:

# A tibble: 2 × 2
     id result
  <dbl>  <dbl>
1     1    1.5
2     2    7  

方式三:分步计算更清晰

先分组计算均值,再关联id计算最终结果,逻辑更直观:

df |> 
  group_by(id) |> 
  summarise(mean_x = mean(x), .groups = "drop") |> 
  mutate(result = mean_x * id) |> 
  select(id, result)

输出同样符合预期。

内容的提问来源于stack exchange,提问作者Mihail

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 03:22:55