如何将dplyr文本处理代码改写为支持传入列参数的通用函数
解决方案
实现思路
- dplyr 1.0+ 版本提供的
across()函数原生支持传入字符向量选择列,配合.names参数可以灵活生成新列名,无需手动实现复杂的tidyeval语法即可完成需求。 - 先将重复的文本处理逻辑抽为独立函数,再通过
across()批量应用到目标列即可。
完整实现代码
library(tidyverse) library(tm) library(glue) library(stringi) # 单条文本处理逻辑 process_single_text <- function(x, stopwords_regex) { x %>% tolower() %>% stringi::stri_trans_general("Latin-ASCII") %>% str_remove_all(stopwords_regex) %>% str_remove_all("[[:punct:]]") %>% str_squish() } # 批量处理文本列的封装函数 process_text_columns <- function(df, columns_list, new_col_names = NULL) { # 预生成停用词正则 stopwords_regex = paste(tm::stopwords('en'), collapse = '\\b|\\b') stopwords_regex = glue('\\b{stopwords_regex}\\b') # 校验新列名长度 if (!is.null(new_col_names) && length(new_col_names) != length(columns_list)) { stop("new_col_names长度需与columns_list保持一致") } # 批量处理列 res <- df %>% mutate( across( .cols = all_of(columns_list), .fns = ~process_single_text(.x, stopwords_regex), .names = "{.col}_proc" ) ) # 若指定新列名 if (!is.null(new_col_names)) { res <- res %>% rename_with(~new_col_names, all_of(paste0(columns_list, "_proc"))) } return(res) }
使用示例
# 构造示例数据 df <- tibble( ocupation = c("Sink Cleaner", "Lion petter"), tasks = c("Cleaning the sink", "Pet the lions"), id = c(1, 2) ) # 方式1:自动生成新列名,列名自动加_proc后缀 df_result1 <- process_text_columns(df, columns_list = c("ocupation", "tasks")) # 方式2:自定义新列名 df_result2 <- process_text_columns(df, columns_list = c("ocupation", "tasks"), new_col_names = c("processed_ocupation", "processed_tasks"))
说明
你需求中提到的sym/!!/!!!等语法是dplyr旧版本实现批量列操作的常用方案,现在有了across()之后大部分批量列操作场景都可以用更简洁的across()实现,不需要手动处理引号、表达式解析等问题,代码可读性和可维护性更高。
内容的提问来源于stack exchange,提问作者Juan C
相关产品推荐
相关产品推荐

