如何实现高效的句子词频统计矩阵生成函数并解决列名设置问题
解决单词出现次数统计矩阵的问题
先修正你现有代码的问题,再提供更高效的实现方案:
一、修正你的原始代码
你的代码存在两个问题:一是矩阵无法直接用rename_with设置列名,二是str_count会匹配子字符串(比如可能误统计包含目标单词的长单词)。修正后的代码如下:
library(stringr) library(purrr) library(dplyr) wordextract <- function(sentences) { # 分割所有句子的单词并去重 words <- unique(unlist(strsplit(sentences, " "))) # 用单词边界精确匹配,避免子串误统计 bag <- map(words, ~ str_count(sentences, regex(paste0("\\b", .x, "\\b"), ignore_case = TRUE))) %>% do.call(cbind, .) # 直接给矩阵设置列名 colnames(bag) <- words # 转成数据框方便查看(可选,保留矩阵也可) as.data.frame(bag) } # 测试示例 sentences <- c("I love bananas", "I hate bananas", "I love apples and I hate bananas") wordextract(sentences)
二、更高效的实现方案
方法1:用tidytext包(简洁易读,适配tidyverse生态)
tidytext专门用于文本数据处理,流程清晰,后续扩展文本预处理(如过滤停用词)也很方便:
library(tidytext) library(dplyr) library(tidyr) word_count_matrix <- function(sentences) { tibble(sentence_id = seq_along(sentences), text = sentences) %>% # 按空格分割单词,保留原大小写 unnest_tokens(word, text, to_lower = FALSE, token = "regex", pattern = " ") %>% # 统计每个句子中各单词的出现次数 count(sentence_id, word) %>% # 转成宽格式,缺失值用0填充 pivot_wider(names_from = word, values_from = n, values_fill = 0) %>% # 移除句子ID列,得到纯计数矩阵 select(-sentence_id) %>% as.matrix() } # 测试 word_count_matrix(sentences)
方法2:用tm包(专业文本挖掘工具,性能更优)
tm包的DocumentTermMatrix是专门生成文档-词频矩阵的工具,适合处理大规模文本:
library(tm) word_count_matrix_tm <- function(sentences) { # 创建文本语料库 corp <- VCorpus(VectorSource(sentences)) # 生成文档-词频矩阵,保留原大小写,按空格分词 dtm <- DocumentTermMatrix(corp, control = list( tokenize = function(x) strsplit(x, " ")[[1]], tolower = FALSE, wordLengths = c(1, Inf) )) # 转换为矩阵格式 as.matrix(dtm) } # 测试 word_count_matrix_tm(sentences)
内容的提问来源于stack exchange,提问作者ManjiroSano
相关产品推荐
相关产品推荐

