dplyr 1.0.7中替代filter_at()与vars(starts_with())的规范语法及动态列过滤实现
在dplyr 1.0.7中替代filter_at()的规范写法
针对你遇到的场景——每个数据框只有一列以identification开头,需要根据该列过滤出匹配ids_to_keep的行,在dplyr 1.0.7里可以直接用filter()结合across()函数来实现,这是替代旧版filter_at()的标准写法:
# 对df_1的过滤 df_1 %>% filter(across(starts_with("identification"), ~ .x %in% ids_to_keep)) # 同样适用于df_2和df_3 df_2 %>% filter(across(starts_with("identification"), ~ .x %in% ids_to_keep)) df_3 %>% filter(across(starts_with("identification"), ~ .x %in% ids_to_keep))
为什么这么写?
在dplyr 1.0.0之后,filter_at()这类"scoped verbs"(作用域动词)被统一替换为across()函数,它可以更灵活地在filter()、mutate()等核心动词中指定要操作的列,语法也更统一直观。
更通用的替代规则
如果你的场景更复杂(比如有多个符合条件的列),可以根据需求选择对应的写法:
- 如果需要任意一列满足条件就保留行,用
if_any()替代旧写法filter_at(vars(...), any_vars(...)) - 如果需要所有列都满足条件才保留行,用
if_all()替代旧写法filter_at(vars(...), all_vars(...))
举个例子,如果你的数据框有多列以identification开头,想要所有这些列都匹配ids_to_keep才保留行,就可以写成:
df %>% filter(if_all(starts_with("identification"), ~ .x %in% ids_to_keep))
而如果是任意一列匹配就保留行,就用:
df %>% filter(if_any(starts_with("identification"), ~ .x %in% ids_to_keep))
这样的写法比旧版的filter_at()更清晰,也更贴合dplyr新版本的设计逻辑。
内容的提问来源于stack exchange,提问作者Emman
相关产品推荐
相关产品推荐

