为何R中grepl统计单词出现次数与手动计数不符?
问题:R语言统计单词出现次数与手动计数不符
我正在开展政客语言使用的相关研究,需要统计文档中某个单词的出现次数,但R语言返回的结果与手动计数不一致。
第一种情况:直接读取XML行统计
以下代码统计单词migrant的出现次数,返回结果为20,但手动打开文档用Ctrl+F查找得到22次:
#Counting the occurrences of the word 'migrant' in a political debate fileContent <- readLines("https://www.theyworkforyou.com/pwdata/scrapedxml/debates/debates2024-01-17c.xml") wordToCount <- c("Migrant") wordCount <- sum(grepl(wordToCount, fileContent, ignore.case = TRUE)) wordCount #returns 20
第二种情况:解析XML后统计
尝试解析XML提取演讲文本后统计,返回结果仅18次,但手动检查导出的output.txt仍能找到22次:
#Same as above but parsing the xml fileContent <- read_xml("https://www.theyworkforyou.com/pwdata/scrapedxml/debates/debates2024-01-17c.xml") fileContent <- xml_find_all(fileContent, ".//speech") fileContent <- xml_text(fileContent) wordToCount <- c("Migrant") wordCount <- sum(grepl(wordToCount, fileContent, ignore.case = TRUE)) wordCount #returns 18 #Outputting the data to double-check values output <- file("output.txt") writeLines(fileContent, output) close(output)
问题原因
两段代码的核心问题都是统计逻辑错误:
grepl()函数的作用是判断单个字符串(这里是每一行/每个演讲片段)是否包含目标词,返回TRUE或FALSE。sum(grepl(...))统计的是包含目标词的行/片段数量,而不是目标词出现的总次数。如果某一行/片段里多次出现migrant,grepl()只会记1次,这就和手动统计的总实例数产生差距。- 第二种情况中,你手动检查导出文本有22次,说明主要问题还是
grepl()的统计逻辑,而非XML节点遗漏。
正确的统计方法
要统计单词出现的总次数,应该使用stringr包的str_count()函数,它会统计每个字符串内目标词的出现次数,再求和:
方法1:直接读取XML行统计
library(stringr) fileContent <- readLines("https://www.theyworkforyou.com/pwdata/scrapedxml/debates/debates2024-01-17c.xml") # 用正则忽略大小写,匹配所有实例 wordToCount <- regex("migrant", ignore_case = TRUE) wordCount <- sum(str_count(fileContent, wordToCount)) wordCount # 应返回22
方法2:解析XML后统计
library(stringr) library(xml2) fileContent <- read_xml("https://www.theyworkforyou.com/pwdata/scrapedxml/debates/debates2024-01-17c.xml") fileContent <- xml_find_all(fileContent, ".//speech") fileContent <- xml_text(fileContent) wordToCount <- regex("migrant", ignore_case = TRUE) wordCount <- sum(str_count(fileContent, wordToCount)) wordCount # 应返回22
额外说明:统计完整单词
如果需要统计独立的完整单词(避免匹配migrants、migrant's这类包含migrant的词),可以给正则加上单词边界符\\b:
wordToCount <- regex("\\bmigrant\\b", ignore_case = TRUE)
内容的提问来源于stack exchange,提问作者C_B
相关产品推荐
相关产品推荐

