Quanteda自定义词典:高频词关联正负词统计与数值赋值技术问询
基于quanteda的词汇关联分析与数值词典问题
数据集与环境
加载所需R包并定义测试文本数据:
library(quanteda) library(quanteda.textstats) df_test<-c("I find water to be so healthy and refreshing", "Nothing like a freshly made burguer to make me feel good", "I dislike sugar in the morning it tastes horrible", "A nice burguer is always crispy and spicy", "It is beyond me to dare to drink soda it's just gross too much sugar", "Yes I will have a hot burguer anytime is so cheap and tasty")
自定义正负情感词典
定义包含正负情感词的词典:
dict_custom <- dictionary(list(positive = c("healthy", "refreshing", "good", "crispy", "spicy", "cheap", "tasty"), negative=c("horrible","gross")))
高频词统计结果
对文本分词后统计Top5高频词:
tok_df<-corpus(df_test) %>% tokens(remove_punct=TRUE) %>% tokens_remove(stopwords("en")) tok_df %>% dfm() %>% textstat_frequency(5)
输出结果:
feature frequency rank docfreq group 1 burguer 3 1 3 all 2 sugar 2 2 2 all 3 find 1 3 1 all 4 water 1 3 1 all 5 healthy 1 3 1 all
问题与解决方案
问题1:获取高频词关联的正负特征词统计及词云
当前代码返回文档级正负计数,若要针对高频词(如burguer)获取其关联文档中的正负特征词出现次数,可按以下步骤实现:
# 筛选包含burguer的文档 corpus_burguer <- corpus_subset(corpus(df_test), grepl("burguer", text)) # 对目标文档分词并预处理 tok_burguer <- tokens(corpus_burguer, remove_punct = TRUE) %>% tokens_remove(stopwords("en")) # 提取词典中的所有情感词,统计其在目标文档中的出现频率 dict_terms <- unlist(dict_custom) freq_burguer <- dfm(tok_burguer) %>% dfm_select(pattern = dict_terms) %>% textstat_frequency() # 为每个情感词添加正负标签 freq_burguer$polarity <- ifelse(freq_burguer$feature %in% dict_custom$positive, "positive", "negative") # 查看统计结果 print(freq_burguer) # 生成情感词云(需先安装wordcloud包) library(wordcloud) wordcloud(freq_burguer$feature, freq_burguer$frequency, color = ifelse(freq_burguer$polarity == "positive", "#2ECC71", "#E74C3C"), scale = c(2, 0.5))
执行后会得到burguer关联文档中各正负情感词的出现次数,词云用不同颜色区分正负词。
问题2:数值词典的合并与数值汇总
若要为词汇分配具体数值而非仅正负标签,可通过命名向量定义数值词典,再匹配到文本数据中进行数值汇总:
# 定义数值词典(命名向量:词为名称,对应数值为值) num_dict <- c(gross = -5, crispy = 5, spicy = 4, tasty = 5, cheap = 3, healthy = 4, refreshing = 3, good = 4, horrible = -4) # 生成完整的文档-特征矩阵 dfm_full <- dfm(tok_df) # 方案1:计算每个文档的情感总分 dfm_num <- dfm_replace(dfm_full, pattern = names(num_dict), replacement = as.character(num_dict)) doc_total_scores <- rowSums(dfm_num, na.rm = TRUE) names(doc_total_scores) <- paste0("text", 1:length(doc_total_scores)) print("各文档情感总分:") print(doc_total_scores) # 方案2:计算每个情感词的总贡献值(词频×对应数值) term_total_scores <- colSums(dfm_full[, names(num_dict)]) * num_dict term_total_scores <- sort(term_total_scores, decreasing = TRUE) print("各情感词总贡献值:") print(term_total_scores)
执行后会分别输出每个文档的情感总分,以及每个情感词的总贡献值(出现次数×对应数值)。
内容的提问来源于stack exchange,提问作者R_Student
相关产品推荐
相关产品推荐

