R语言词典法情感分析按文章长度标准化计算报错咨询
词典法新闻情感分析标准化得分计算报错修复
场景说明
- 分析对象为新闻报道,分析单元为单篇报道,不做句子级拆分
- 回归分析前需按单篇文章长度对情感词得分做标准化处理
原始代码
# 数据清洗与分词 data.C <- read.csv("filelocation",stringsAsFactors = FALSE) data.C <- data.frame(data.C) data.C <- data.C %>% unnest_tokens(word, sentence) data.C <- data.C %>% anti_join(stop_words, by= "word") # 匹配情感词典计算得分 struggle <- data.frame(data.C) %>% inner_join(get_sentiments("afinn")) struggle %>% group_by(story_index) %>% summarize(weights(sum(value)/n()))
触发报错
Error in `summarize()`: ! Problem while computing `..1 = weights(sum(value)/n())`. i The error occurred in group 1: story_index = 1. Caused by error: ! $ operator is invalid for atomic vectors
问题原因
- 语法错误:
summarize()中新列赋值缺少等号,weights(sum(value)/n())会被识别为调用weights()函数传入计算结果,而非创建名为weights的结果列。R内置weights()方法不支持原子向量输入,因此触发报错。 - 逻辑偏差:当前分组内的
n()统计的是匹配到AFINN情感词典的词数,不是单篇报道去停用词后的总词数,直接作为分母不符合“按文章长度标准化”的要求,会导致得分估计偏差。
修正后实现代码
library(tidyverse) library(tidytext) # 1. 读取数据、分词、去停用词 data.C <- read.csv("filelocation", stringsAsFactors = FALSE) %>% unnest_tokens(word, sentence) %>% anti_join(stop_words, by = "word") # 2. 统计每篇报道去停用词后的总有效词数(即文章长度基准) doc_total_words <- data.C %>% group_by(story_index) %>% summarise(total_valid_words = n(), .groups = "drop") # 3. 匹配情感词典,计算长度标准化后的情感得分 sentiment_std_result <- data.C %>% inner_join(get_sentiments("afinn"), by = "word") %>% group_by(story_index) %>% summarise(sentiment_sum = sum(value), .groups = "drop") %>% left_join(doc_total_words, by = "story_index") %>% # 标准化得分 = 单篇情感得分总和 / 单篇总有效词数 mutate(std_sentiment_score = sentiment_sum / total_valid_words)
计算得到的std_sentiment_score字段即为按文章长度标准化后的单篇报道情感得分,可直接用于后续回归分析。
内容的提问来源于stack exchange,提问作者meriem
相关产品推荐
相关产品推荐

