使用R进行文本挖掘:如何查看文档中的正负情感词?
查看具体情感词汇的解决方案
嘿,作为R语言新手能走到这一步已经很棒啦!你已经算出了正负情感词的数量,现在要查看具体的词汇其实只需要在现有代码上稍作调整就行,我给你几种实用的方案:
方案1:查看所有匹配到的情感词(含重复)
这个会保留文本中所有出现过的情感词,包括重复出现的,能让你看到每个词的实际出现情况:
library(readr) library(tidyverse) library(tidytext) library(glue) library(stringr) library(dplyr) # 你的原有文本处理代码 davos <- read_file("davos.txt") fileText <- glue(read_file(davos)) fileText <- gsub("\\$", "", fileText) tokens <- data_frame(text = fileText) %>% unnest_tokens(word, text) # 获取所有情感词(包含重复出现的) all_sentiment_words <- tokens %>% inner_join(get_sentiments("bing")) %>% select(word, sentiment) # 只保留词汇和对应的情感类型 # 预览前几行结果 head(all_sentiment_words) # 查看全部结果 print(all_sentiment_words)
方案2:查看去重后的独特情感词
如果你只想知道文本里出现过哪些不同的正负情感词,不需要重复项,就用这个:
# 在分词基础上获取去重后的情感词 unique_sentiment_words <- tokens %>% inner_join(get_sentiments("bing")) %>% distinct(word, sentiment) # 去重,保留每个词汇+情感的唯一组合 # 分别提取正面和负面词 positive_unique <- unique_sentiment_words %>% filter(sentiment == "positive") negative_unique <- unique_sentiment_words %>% filter(sentiment == "negative") # 查看结果 cat("正面情感词列表:\n") print(positive_unique) cat("\n负面情感词列表:\n") print(negative_unique)
方案3:查看带出现次数的情感词(最实用)
这个方案不仅能看到具体词汇,还能知道每个词在文本里的出现频次,按频次排序后更直观:
sentiment_word_counts <- tokens %>% inner_join(get_sentiments("bing")) %>% count(word, sentiment, sort = TRUE) # 按词汇和情感分组计数,降序排序 # 查看结果 print(sentiment_word_counts)
你可以根据自己的需求选其中一种,比如想了解高频情感词就选方案3,只想看有哪些独特词汇就选方案2~
内容的提问来源于stack exchange,提问作者Emrah BILGIC
相关产品推荐
相关产品推荐

