如何基于词汇占比过滤情感模型识别不足的Reddit帖子?
问题背景
我有一个包含40K条Reddit帖子的数据集,正在使用含8K个独特词汇及短语的词典式模型进行单帖情感估算。此前我通过固定数量规则过滤帖子(剔除恰好识别出1个正面词+1个负面词的帖子),代码如下:
#Loading packages library(tidyverse) require(readxl) require(writexl) library(quanteda) library(stm) library(stmCorrViz) library(stringi)
#Filtering out posts where ONLY 1 positive and 1 negative words are recognized ## valences_by_post_oneword <- valences_by_post # subset data for posts where only one word was recognized one_pn<- valences_by_post_oneword %>% filter(positive==1 & negative==1) valences_by_post_oneword <- valences_by_post_oneword %>% filter(!(positive==1 & negative==1)) #2011 & 2012 valence_oneowrd<-valences_by_post_oneword %>% filter(year == 2011 | year ==2012)%>% group_by(month_year) %>% summarize(mean_valence= mean(valence), n=n())
考虑到帖子总词汇量存在差异,采用占比过滤更为合理,我需要过滤掉词典识别的情感词汇总数占该帖总词汇量5%及以下的帖子。当前数据集结构如下:
| post | negative_words | positive_words | total_words | valence_score |
|---|---|---|---|---|
| xyz. | 2 | 1 | 10 | -0.66 |
其中negative_words、positive_words是词典识别的情感词汇数量,total_words是帖子总词汇量,valence_score按公式(negative_words + positive_words)/total_words计算。
解决方案代码
可以通过新增占比列再过滤的方式实现需求,代码如下:
# 加载依赖包(与原有代码一致) library(tidyverse) require(readxl) require(writexl) library(quanteda) library(stm) library(stmCorrViz) library(stringi) # 计算情感词汇占比并过滤掉占比≤5%的帖子 valences_filtered <- valences_by_post %>% # 新增列:情感词汇占总词汇量的比例 mutate(sentiment_word_ratio = (negative_words + positive_words)/total_words) %>% # 过滤掉占比≤5%的帖子,保留占比超过5%的 filter(sentiment_word_ratio > 0.05) # 针对2011、2012年数据执行分组汇总(对齐原有逻辑) valence_filtered_2011_2012 <- valences_filtered %>% filter(year %in% c(2011, 2012)) %>% group_by(month_year) %>% summarize(mean_valence = mean(valence_score), n = n())
代码说明
mutate语句新增sentiment_word_ratio列,统一计算情感词汇占总词汇量的比例,避免重复计算;filter(sentiment_word_ratio > 0.05)直接剔除占比≤5%的帖子,保留符合要求的样本;- 后续分组汇总逻辑与原有代码一致,注意将原有代码中的
valence替换为数据结构中的valence_score,避免变量名错误。
内容的提问来源于stack exchange,提问作者nesta1990
相关产品推荐
相关产品推荐

