dplyr 1.0.2中scoped summarise与n()的正确语法咨询
Hey there, let's work through your problem with scoped summarise functions in dplyr 1.0.2. You're trying to replicate the behavior of your standard summarise call using scoped verbs like summarise_if, but your attempted runs are throwing errors—let's break down why, then fix it.
First, Why Your Attempts Failed
- Attempt 1 Error: The message about
n()makes perfect sense here.summarise_ifexpects functions that operate on individual columns, butn()is a group-level function (it counts rows per group, not values in a single column). Also, your syntax forsum()is off—you need to use a formula (~sum(.x)) to reference the column being processed instead of callingsum()directly. - Attempt 2 Error:
summarise_ifonly accepts one function (or a named list of functions) as its second argument. Passing multiple~expressions like you did doesn't fit its syntax rules.
Correct Implementations
Option 1: Proper summarise_if Syntax
Since you want two types of calculations—column-specific sums for numeric columns, plus group-level row counts and proportions—we can structure this to handle both:
mtcars %>% group_by(am, gear) %>% # Calculate sum for all double columns first summarise_if(is.double, list(sum = ~sum(.x)), .groups = "keep") %>% # Add group-level n and proportion afterward mutate(n = n(), prop = sum_disp / n)
Or, if you prefer to do it all within summarise_if, you can add the group-level metrics as separate arguments outside the column function list:
mtcars %>% group_by(am, gear) %>% summarise_if( is.double, list(sum = ~sum(.x)), # Add group-specific metrics here n = n(), prop = sum(disp) / n(), .groups = "keep" )
Note: The .groups = "keep" parameter is new in dplyr 1.0+ to preserve your grouping structure—omit it if you don't need to keep grouping after summarising.
Option 2: Use across() (Recommended for dplyr 1.0+)
Scoped verbs like summarise_if are being phased out in favor of across() in dplyr 1.0 and later. It's more flexible and easier to read. Here's how to replicate your original logic with across():
If you only need calculations for the disp column:
mtcars %>% group_by(am, gear) %>% summarise( sum = sum(disp), n = n(), prop = sum / n, .groups = "keep" )
If you want to apply the sum function to all numeric columns and include the group metrics:
mtcars %>% group_by(am, gear) %>% summarise( across(is.double, list(sum = ~sum(.x))), n = n(), prop = sum_disp / n, .groups = "keep" )
Verify the Output
Either approach will give you the exact same result as your original non-scoped code:
# A tibble: 4 × 5 # Groups: am, gear [4] am gear sum n prop <dbl> <dbl> <dbl> <int> <dbl> 1 0 3 4614 15 308. 2 0 4 1166 4 291. 3 1 4 1156 8 145. 4 1 5 729 5 146.
内容的提问来源于stack exchange,提问作者Paul

