You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

dplyr 1.0.2中scoped summarise与n()的正确语法咨询

Fixing Scoped Summarise Syntax Issues in dplyr 1.0.2

Hey there, let's work through your problem with scoped summarise functions in dplyr 1.0.2. You're trying to replicate the behavior of your standard summarise call using scoped verbs like summarise_if, but your attempted runs are throwing errors—let's break down why, then fix it.

First, Why Your Attempts Failed

  • Attempt 1 Error: The message about n() makes perfect sense here. summarise_if expects functions that operate on individual columns, but n() is a group-level function (it counts rows per group, not values in a single column). Also, your syntax for sum() is off—you need to use a formula (~sum(.x)) to reference the column being processed instead of calling sum() directly.
  • Attempt 2 Error: summarise_if only accepts one function (or a named list of functions) as its second argument. Passing multiple ~ expressions like you did doesn't fit its syntax rules.

Correct Implementations

Option 1: Proper summarise_if Syntax

Since you want two types of calculations—column-specific sums for numeric columns, plus group-level row counts and proportions—we can structure this to handle both:

mtcars %>% 
  group_by(am, gear) %>% 
  # Calculate sum for all double columns first
  summarise_if(is.double, list(sum = ~sum(.x)), .groups = "keep") %>% 
  # Add group-level n and proportion afterward
  mutate(n = n(), prop = sum_disp / n)

Or, if you prefer to do it all within summarise_if, you can add the group-level metrics as separate arguments outside the column function list:

mtcars %>% 
  group_by(am, gear) %>% 
  summarise_if(
    is.double, 
    list(sum = ~sum(.x)),
    # Add group-specific metrics here
    n = n(),
    prop = sum(disp) / n(),
    .groups = "keep"
  )

Note: The .groups = "keep" parameter is new in dplyr 1.0+ to preserve your grouping structure—omit it if you don't need to keep grouping after summarising.

Scoped verbs like summarise_if are being phased out in favor of across() in dplyr 1.0 and later. It's more flexible and easier to read. Here's how to replicate your original logic with across():

If you only need calculations for the disp column:

mtcars %>% 
  group_by(am, gear) %>% 
  summarise(
    sum = sum(disp),
    n = n(),
    prop = sum / n,
    .groups = "keep"
  )

If you want to apply the sum function to all numeric columns and include the group metrics:

mtcars %>% 
  group_by(am, gear) %>% 
  summarise(
    across(is.double, list(sum = ~sum(.x))),
    n = n(),
    prop = sum_disp / n,
    .groups = "keep"
  )

Verify the Output

Either approach will give you the exact same result as your original non-scoped code:

# A tibble: 4 × 5
# Groups:   am, gear [4]
     am  gear   sum     n  prop
  <dbl> <dbl> <dbl> <int> <dbl>
1     0     3  4614    15  308.
2     0     4  1166     4  291.
3     1     4  1156     8  145.
4     1     5   729     5  146.

内容的提问来源于stack exchange,提问作者Paul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:38:55