如何在dplyr的summarise中使用行值,实现类似data.table的效果?
解决dplyr分组计算时重复行与警告的正确方式
先看示例数据:
library(dplyr) library(data.table) df <- data.frame( id = c(1,1,2,2), x = 1:4 )
需求很明确:按id分组,计算每组x的均值后乘以该组的id值。
用data.table实现起来很直观:
setDT(df)[, .(mean(x) * id), by=id]
输出结果符合预期,每组一行:
id V1 1: 1 1.5 2: 2 7.0
但用dplyr的summarise直接写时,会出现重复行和警告:
df |> group_by(id) |> summarise(mean(x) * id)
输出及警告:
`summarise()` has grouped output by 'id'. You can override using the `.groups` argument. # A tibble: 4 × 2 # Groups: id [2] id `mean(x) * id` <dbl> <dbl> 1 1 1.5 2 1 1.5 3 2 7 4 2 7 Warning message: Returning more (or less) than 1 row per `summarise()` group was deprecated in dplyr 1.1.0. ℹ Please use `reframe()` instead.
虽然加unique()能去重,但这不是dplyr的正确用法,这里提供几种规范的解决方式:
方式一:用reframe()替代summarise()
dplyr 1.1.0之后推荐用reframe()处理非单行的分组结果,这里我们可以明确取组内的id值(因为组内id都相同,用first(id)即可):
df |> group_by(id) |> reframe(result = mean(x) * first(id))
输出:
# A tibble: 2 × 2 id result <dbl> <dbl> 1 1 1.5 2 2 7
方式二:在summarise()中确保结果为标量
summarise()要求每组返回单行结果,所以需要把id转换成标量(组内id唯一,用first(id)或unique(id)都可以),同时可以加上.groups = "drop"取消分组状态:
df |> group_by(id) |> summarise(result = mean(x) * first(id), .groups = "drop")
输出:
# A tibble: 2 × 2 id result <dbl> <dbl> 1 1 1.5 2 2 7
方式三:分步计算更清晰
先分组计算均值,再关联id计算最终结果,逻辑更直观:
df |> group_by(id) |> summarise(mean_x = mean(x), .groups = "drop") |> mutate(result = mean_x * id) |> select(id, result)
输出同样符合预期。
内容的提问来源于stack exchange,提问作者Mihail
相关产品推荐
相关产品推荐

