如何在R的dplyr中用across函数筛选未被其他across捕获的剩余列?
在dplyr中使用across选择未被其他语句捕获的剩余列的通用实现
需求场景
在使用dplyr进行数据汇总时,需要实现以下固定逻辑:
- 对所有字符型字段应用自定义函数
first.if.unique(组内值唯一时取第一个值,否则返回NA) - 对指定数值列(如
mass)应用mean函数求均值 - 对剩余所有未被前两步处理的列应用
sum函数求和
由于列数量多且选择规则灵活,手动枚举剩余列效率极低,需要通用的实现方案。
通用解决方案
利用tidyselect的选择语法,在最后一个across中通过排除已处理列的方式自动选中剩余列,无需手动维护剩余列列表,完全适配列结构变化的场景:
- 用
where(is.character)选中所有字符列处理 - 用
any_of(cols_to_average)选中指定列处理 - 用
-c(where(is.character), all_of(cols_to_average))排除前两步的列,自动选中剩余列应用sum
示例代码
library(dplyr) # 自定义函数:组内值唯一则取第一个,否则返回NA first.if.unique = function(x) if(length(unique(x)) == 1) x[1] else NA # 构造示例数据集 x = starwars %>% select(species, sex, mass, height) %>% head(10) # 指定需要求均值的列 cols_to_average = c("mass") # 通用汇总逻辑 x %>% group_by(sex) %>% summarise( # 处理字符型列 across(where(is.character), first.if.unique), # 处理指定均值列(添加na.rm避免NA值干扰计算) across(any_of(cols_to_average), ~mean(.x, na.rm = TRUE)), # 处理剩余所有列(排除已处理的字符列和均值列) across(-c(where(is.character), all_of(cols_to_average)), sum), # 取消分组,返回普通tibble .groups = "drop" )
输入数据集
print(x) #> # A tibble: 10 × 4 #> species sex mass height #> <chr> <chr> <dbl> <int> #> 1 Human male 77 172 #> 2 Droid none 75 167 #> 3 Droid none 32 96 #> 4 Human male 136 202 #> 5 Human female 49 150 #> 6 Human male 120 178 #> 7 Human female 75 165 #> 8 Droid none 32 97 #> 9 Human male 84 183 #> 10 Human male 77 182
预期输出
#> # A tibble: 3 × 4 #> sex species mass height #> <chr> <chr> <dbl> <int> #> 1 female Human 62 315 #> 2 male Human 98.8 917 #> 3 none Droid 46.3 360
内容的提问来源于stack exchange,提问作者user425678
相关产品推荐
相关产品推荐

