如何在dplyr工作流中统计str_replace_all的替换次数?
高效实现dplyr工作流中字符串替换的次数统计
在使用dplyr结合str_replace_all清理数据框时,我们需要同时生成包含每个拼写错误替换次数的汇总表,且要保证处理大型数据的效率。以下是实现方案:
核心思路
通过一次遍历原始数据完成替换操作,同时利用向量化的字符串计数函数统计每个拼写错误的总出现次数(即替换次数),避免重复扫描数据消耗时间。
完整代码实现
library(dplyr) library(stringr) library(tibble) library(flextable) # 可选:大数据量时用stringi提升速度 # library(stringi) # 1. 定义原始数据与修正规则 my_str <- c('a', 'ca', 'c', 'bla') my_df <- data.frame(my_str) my_typo_corrections <- c( 'a' = 'b', 'c' = 'd', 'what so ever' = 'whatever' ) # 2. 执行字符串替换,得到修正后的数据框 my_corrected_df <- my_df %>% mutate(my_str_corrected = str_replace_all(my_str, my_typo_corrections)) # 3. 统计每个拼写错误的总替换次数 typo_counts <- purrr::map_dfr(names(my_typo_corrections), function(typo) { # 用str_count(或stri_count_fixed)统计全局匹配次数 total_count <- sum(str_count(my_df$my_str, fixed(typo))) # 大数据量推荐替换为:sum(stri_count_fixed(my_df$my_str, typo)) tibble( Typo = typo, Correction = my_typo_corrections[typo], Replace_Count = total_count ) }) # 4. 生成带替换次数的汇总表 flextable(typo_counts) %>% autofit()
关键细节说明
- 使用
fixed(typo)确保按字面匹配字符串,避免正则表达式的转义问题,同时提升匹配效率 - 若处理超大型数据,推荐用
stringi::stri_count_fixed替代str_count,底层C实现的速度远快于基础stringr函数 purrr::map_dfr将每个拼写错误的统计结果合并为数据框,比循环更高效且代码更简洁
内容的提问来源于stack exchange,提问作者MsGISRocker
相关产品推荐
相关产品推荐

