You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在dplyr工作流中统计str_replace_all的替换次数?

高效实现dplyr工作流中字符串替换的次数统计

在使用dplyr结合str_replace_all清理数据框时,我们需要同时生成包含每个拼写错误替换次数的汇总表,且要保证处理大型数据的效率。以下是实现方案:

核心思路

通过一次遍历原始数据完成替换操作,同时利用向量化的字符串计数函数统计每个拼写错误的总出现次数(即替换次数),避免重复扫描数据消耗时间。

完整代码实现

library(dplyr)
library(stringr)
library(tibble)
library(flextable)
# 可选:大数据量时用stringi提升速度
# library(stringi)

# 1. 定义原始数据与修正规则
my_str <- c('a', 'ca', 'c', 'bla')
my_df <- data.frame(my_str)

my_typo_corrections <- c(
  'a' = 'b',
  'c' = 'd',
  'what so ever' = 'whatever'
)

# 2. 执行字符串替换,得到修正后的数据框
my_corrected_df <- my_df %>%
  mutate(my_str_corrected = str_replace_all(my_str, my_typo_corrections))

# 3. 统计每个拼写错误的总替换次数
typo_counts <- purrr::map_dfr(names(my_typo_corrections), function(typo) {
  # 用str_count(或stri_count_fixed)统计全局匹配次数
  total_count <- sum(str_count(my_df$my_str, fixed(typo)))
  # 大数据量推荐替换为:sum(stri_count_fixed(my_df$my_str, typo))
  
  tibble(
    Typo = typo,
    Correction = my_typo_corrections[typo],
    Replace_Count = total_count
  )
})

# 4. 生成带替换次数的汇总表
flextable(typo_counts) %>%
  autofit()

关键细节说明

  • 使用fixed(typo)确保按字面匹配字符串,避免正则表达式的转义问题,同时提升匹配效率
  • 若处理超大型数据,推荐用stringi::stri_count_fixed替代str_count,底层C实现的速度远快于基础stringr函数
  • purrr::map_dfr将每个拼写错误的统计结果合并为数据框,比循环更高效且代码更简洁

内容的提问来源于stack exchange,提问作者MsGISRocker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 19:23:14