You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现高效的句子词频统计矩阵生成函数并解决列名设置问题

解决单词出现次数统计矩阵的问题

先修正你现有代码的问题,再提供更高效的实现方案:

一、修正你的原始代码

你的代码存在两个问题:一是矩阵无法直接用rename_with设置列名,二是str_count会匹配子字符串(比如可能误统计包含目标单词的长单词)。修正后的代码如下:

library(stringr)
library(purrr)
library(dplyr)

wordextract <- function(sentences) {
  # 分割所有句子的单词并去重
  words <- unique(unlist(strsplit(sentences, " ")))
  # 用单词边界精确匹配,避免子串误统计
  bag <- map(words, ~ str_count(sentences, regex(paste0("\\b", .x, "\\b"), ignore_case = TRUE))) %>%
    do.call(cbind, .)
  # 直接给矩阵设置列名
  colnames(bag) <- words
  # 转成数据框方便查看(可选,保留矩阵也可)
  as.data.frame(bag)
}

# 测试示例
sentences <- c("I love bananas", "I hate bananas", "I love apples and I hate bananas")
wordextract(sentences)

二、更高效的实现方案

方法1:用tidytext包(简洁易读,适配tidyverse生态)

tidytext专门用于文本数据处理,流程清晰,后续扩展文本预处理(如过滤停用词)也很方便:

library(tidytext)
library(dplyr)
library(tidyr)

word_count_matrix <- function(sentences) {
  tibble(sentence_id = seq_along(sentences), text = sentences) %>%
    # 按空格分割单词,保留原大小写
    unnest_tokens(word, text, to_lower = FALSE, token = "regex", pattern = " ") %>%
    # 统计每个句子中各单词的出现次数
    count(sentence_id, word) %>%
    # 转成宽格式,缺失值用0填充
    pivot_wider(names_from = word, values_from = n, values_fill = 0) %>%
    # 移除句子ID列,得到纯计数矩阵
    select(-sentence_id) %>%
    as.matrix()
}

# 测试
word_count_matrix(sentences)

方法2:用tm包(专业文本挖掘工具,性能更优)

tm包的DocumentTermMatrix是专门生成文档-词频矩阵的工具,适合处理大规模文本:

library(tm)

word_count_matrix_tm <- function(sentences) {
  # 创建文本语料库
  corp <- VCorpus(VectorSource(sentences))
  # 生成文档-词频矩阵,保留原大小写,按空格分词
  dtm <- DocumentTermMatrix(corp, control = list(
    tokenize = function(x) strsplit(x, " ")[[1]],
    tolower = FALSE,
    wordLengths = c(1, Inf)
  ))
  # 转换为矩阵格式
  as.matrix(dtm)
}

# 测试
word_count_matrix_tm(sentences)

内容的提问来源于stack exchange,提问作者ManjiroSano

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 16:45:47