You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R语言tm包在Corpus中统计指定词表的文档词频?

用tm包统计指定词表的文档词频

前提

假设你已经加载了tm包,且你的语料库对象名为my_corpus,指定词表为target_words(比如月份列表)。


方法一:基于词-文档矩阵(TermDocumentMatrix)

这是tm包的标准方法,适合快速批量统计:

  1. 预处理语料库(可选但推荐)
    先统一文本格式,避免大小写、标点等导致的匹配误差:
library(tm)
# 转小写
my_corpus <- tm_map(my_corpus, content_transformer(tolower))
# 按需去除标点、数字(可选)
my_corpus <- tm_map(my_corpus, removePunctuation)
my_corpus <- tm_map(my_corpus, removeNumbers)
  1. 构建词-文档矩阵
tdm <- TermDocumentMatrix(my_corpus)
  1. 提取目标词的频次数据
# 定义你的目标词表
target_words <- c("january", "february", "march", "april", "may", "june", 
                  "july", "august", "september", "october", "november", "december")

# 筛选语料库中实际存在的目标词(避免不存在的词引发报错)
existing_words <- intersect(target_words, rownames(tdm))

# 提取频次并转置,得到「文档×目标词」的结构
word_freq <- t(as.matrix(tdm[existing_words, ]))

# 补充未出现词的0值,并按原词表顺序排列列
missing_words <- setdiff(target_words, existing_words)
if (length(missing_words) > 0) {
  word_freq <- cbind(word_freq, matrix(0, nrow = nrow(word_freq), ncol = length(missing_words),
                                       dimnames = list(NULL, missing_words)))
  word_freq <- word_freq[, target_words]
}
  1. 查看结果
    word_freq即为所需数据,每行对应一份文档,每列是对应目标词的出现频次:
head(word_freq)

方法二:逐个文档遍历统计

如果需要自定义分词逻辑,这种方法更灵活:

# 定义统计函数:输入单个文档,返回目标词的频次向量
count_target_words <- function(doc) {
  text_content <- content(doc)
  # 拆分文本为单词(可替换为更严谨的分词逻辑,比如用stringr包)
  word_list <- strsplit(text_content, "\\s+")[[1]]
  # 统计每个目标词的出现次数
  sapply(target_words, function(word) sum(word_list == word))
}

# 应用函数到整个语料库
word_freq_list <- tm_map(my_corpus, count_target_words)
# 转换为数据框格式
word_freq_df <- do.call(rbind, word_freq_list)
# 设置行名为文档ID(与原语料库对应)
rownames(word_freq_df) <- names(my_corpus)

内容的提问来源于stack exchange,提问作者Li4991

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 20:04:49