You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何R中grepl统计单词出现次数与手动计数不符?

问题:R语言统计单词出现次数与手动计数不符

我正在开展政客语言使用的相关研究,需要统计文档中某个单词的出现次数,但R语言返回的结果与手动计数不一致。

第一种情况:直接读取XML行统计

以下代码统计单词migrant的出现次数,返回结果为20,但手动打开文档用Ctrl+F查找得到22次:

#Counting the occurrences of the word 'migrant' in a political debate
fileContent <- readLines("https://www.theyworkforyou.com/pwdata/scrapedxml/debates/debates2024-01-17c.xml")
wordToCount <- c("Migrant") 
wordCount <- sum(grepl(wordToCount, fileContent, ignore.case = TRUE))
wordCount #returns 20

第二种情况:解析XML后统计

尝试解析XML提取演讲文本后统计,返回结果仅18次,但手动检查导出的output.txt仍能找到22次:

#Same as above but parsing the xml
fileContent <- read_xml("https://www.theyworkforyou.com/pwdata/scrapedxml/debates/debates2024-01-17c.xml")
fileContent <- xml_find_all(fileContent, ".//speech")
fileContent <- xml_text(fileContent)
wordToCount <- c("Migrant") 
wordCount <- sum(grepl(wordToCount, fileContent, ignore.case = TRUE))
wordCount #returns 18

#Outputting the data to double-check values
output <- file("output.txt")
writeLines(fileContent, output)
close(output)

问题原因

两段代码的核心问题都是统计逻辑错误:

  • grepl()函数的作用是判断单个字符串(这里是每一行/每个演讲片段)是否包含目标词,返回TRUE或FALSE。sum(grepl(...))统计的是包含目标词的行/片段数量,而不是目标词出现的总次数。如果某一行/片段里多次出现migrant,grepl()只会记1次,这就和手动统计的总实例数产生差距。
  • 第二种情况中,你手动检查导出文本有22次,说明主要问题还是grepl()的统计逻辑,而非XML节点遗漏。

正确的统计方法

要统计单词出现的总次数,应该使用stringr包的str_count()函数,它会统计每个字符串内目标词的出现次数,再求和:

方法1:直接读取XML行统计

library(stringr)
fileContent <- readLines("https://www.theyworkforyou.com/pwdata/scrapedxml/debates/debates2024-01-17c.xml")
# 用正则忽略大小写,匹配所有实例
wordToCount <- regex("migrant", ignore_case = TRUE)
wordCount <- sum(str_count(fileContent, wordToCount))
wordCount # 应返回22

方法2:解析XML后统计

library(stringr)
library(xml2)
fileContent <- read_xml("https://www.theyworkforyou.com/pwdata/scrapedxml/debates/debates2024-01-17c.xml")
fileContent <- xml_find_all(fileContent, ".//speech")
fileContent <- xml_text(fileContent)
wordToCount <- regex("migrant", ignore_case = TRUE)
wordCount <- sum(str_count(fileContent, wordToCount))
wordCount # 应返回22

额外说明:统计完整单词

如果需要统计独立的完整单词(避免匹配migrants、migrant's这类包含migrant的词),可以给正则加上单词边界符\\b:

wordToCount <- regex("\\bmigrant\\b", ignore_case = TRUE)

内容的提问来源于stack exchange,提问作者C_B

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 11:00:08