如何在R中统计向量词汇在字符串中的出现总次数(类preg_match_all)
问题:统计向量词汇在字符串中的出现总频次(类似PHP preg_match_all的R实现)
背景与需求
给定文本字符串和词汇向量:
String: "Auch ein blindes Huhn findet einmal ein Korn." Vector: "auch", "ein"
需要统计向量中每个词汇的出现次数,最终总频次应为3。
当前尝试的局限
目前仅能判断词汇是否存在并统计存在的词汇数量,代码如下:
library(stringr) deu <- c("\\bauch\\b", "\\bein\\b") str_detect(tolower("Auch ein blindes Huhn findet einmal ein Korn."), deu) [1] TRUE TRUE sum(str_detect(tolower("Auch ein blindes Huhn findet einmal ein Korn."), deu)) [1] 2
但str_detect仅返回词汇是否存在(TRUE/FALSE),而非实际出现次数(1,2),求和结果不符合需求。
希望找到R中类似PHP preg_match_all的函数,PHP示例代码如下:
preg_match_all("/\bauch\b|\bein\b/i", "Auch ein blindes Huhn findet einmal ein Korn.", $matches); print_r($matches); Array ( [0] => Array ( [0] => Auch [1] => ein [2] => ein ) ) echo preg_match_all("/\bauch\b|\bein\b/i", "Auch ein blindes Huhn findet einmal ein Korn.", $matches); 3
要求避免使用循环,且已查阅过类似问题,但要么未统计出现次数,要么未使用模式向量进行搜索,若标记此问题为重复,请确保重复问题与本问题完全一致。
解决方案
方法1:使用stringr::str_count直接统计
str_count支持对模式向量逐个统计出现次数,求和即可得到总频次:
library(stringr) text <- "Auch ein blindes Huhn findet einmal ein Korn." patterns <- c("\\bauch\\b", "\\bein\\b") # 统一转为小写匹配,消除大小写差异 counts_per_word <- str_count(tolower(text), patterns) # counts_per_word 结果为 [1] 1 2 total_count <- sum(counts_per_word) # total_count 结果为 3
方法2:合并模式后用str_match_all一次性匹配
将所有目标词汇合并为一个正则表达式,用str_match_all获取所有匹配结果,再统计结果长度:
library(stringr) text <- "Auch ein blindes Huhn findet einmal ein Korn." target_words <- c("auch", "ein") # 构建包含单词边界、不区分大小写的合并正则 combined_regex <- str_c("\\b(", str_c(target_words, collapse = "|"), ")\\b", ignore_case = TRUE) # 获取所有匹配项 all_matches <- str_match_all(text, combined_regex) # 统计总匹配数 total_count <- length(unlist(all_matches)) # total_count 结果为 3
方法3:使用base R的gregexpr(无需额外包)
用base R原生函数gregexpr统计每个模式的出现次数,再求和:
text <- "Auch ein blindes Huhn findet einmal ein Korn." patterns <- c("\\bauch\\b", "\\bein\\b") # 对每个模式统计出现次数 counts_per_word <- sapply(patterns, function(pattern) { match_positions <- gregexpr(pattern, tolower(text), fixed = FALSE)[[1]] sum(match_positions != -1) }) # counts_per_word 结果为 auch ein # 1 2 total_count <- sum(counts_per_word) # total_count 结果为 3
内容的提问来源于stack exchange,提问作者Ben
相关产品推荐
相关产品推荐

