You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中统计向量词汇在字符串中的出现总次数(类preg_match_all)

问题:统计向量词汇在字符串中的出现总频次(类似PHP preg_match_all的R实现)

背景与需求

给定文本字符串和词汇向量:

String: "Auch ein blindes Huhn findet einmal ein Korn."
Vector: "auch", "ein"

需要统计向量中每个词汇的出现次数,最终总频次应为3。

当前尝试的局限

目前仅能判断词汇是否存在并统计存在的词汇数量,代码如下:

library(stringr)
deu <- c("\\bauch\\b", "\\bein\\b")
str_detect(tolower("Auch ein blindes Huhn findet einmal ein Korn."), deu)

[1] TRUE TRUE

sum(str_detect(tolower("Auch ein blindes Huhn findet einmal ein Korn."), deu))

[1] 2

但str_detect仅返回词汇是否存在(TRUE/FALSE),而非实际出现次数(1,2),求和结果不符合需求。

希望找到R中类似PHP preg_match_all的函数,PHP示例代码如下:

preg_match_all("/\bauch\b|\bein\b/i", "Auch ein blindes Huhn findet einmal ein Korn.", $matches);
print_r($matches);

Array
(
    [0] => Array
        (
            [0] => Auch
            [1] => ein
            [2] => ein
        )

)

echo preg_match_all("/\bauch\b|\bein\b/i", "Auch ein blindes Huhn findet einmal ein Korn.", $matches);

3

要求避免使用循环,且已查阅过类似问题,但要么未统计出现次数,要么未使用模式向量进行搜索,若标记此问题为重复,请确保重复问题与本问题完全一致。


解决方案

方法1:使用stringr::str_count直接统计

str_count支持对模式向量逐个统计出现次数,求和即可得到总频次:

library(stringr)

text <- "Auch ein blindes Huhn findet einmal ein Korn."
patterns <- c("\\bauch\\b", "\\bein\\b")

# 统一转为小写匹配,消除大小写差异
counts_per_word <- str_count(tolower(text), patterns)
# counts_per_word 结果为 [1] 1 2
total_count <- sum(counts_per_word)
# total_count 结果为 3

方法2:合并模式后用str_match_all一次性匹配

将所有目标词汇合并为一个正则表达式,用str_match_all获取所有匹配结果,再统计结果长度:

library(stringr)

text <- "Auch ein blindes Huhn findet einmal ein Korn."
target_words <- c("auch", "ein")

# 构建包含单词边界、不区分大小写的合并正则
combined_regex <- str_c("\\b(", str_c(target_words, collapse = "|"), ")\\b", ignore_case = TRUE)

# 获取所有匹配项
all_matches <- str_match_all(text, combined_regex)
# 统计总匹配数
total_count <- length(unlist(all_matches))
# total_count 结果为 3

方法3:使用base R的gregexpr(无需额外包)

用base R原生函数gregexpr统计每个模式的出现次数,再求和:

text <- "Auch ein blindes Huhn findet einmal ein Korn."
patterns <- c("\\bauch\\b", "\\bein\\b")

# 对每个模式统计出现次数
counts_per_word <- sapply(patterns, function(pattern) {
  match_positions <- gregexpr(pattern, tolower(text), fixed = FALSE)[[1]]
  sum(match_positions != -1)
})
# counts_per_word 结果为 auch ein 
#                1   2
total_count <- sum(counts_per_word)
# total_count 结果为 3

内容的提问来源于stack exchange,提问作者Ben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 07:04:54