如何从段落中提取积极/中立/消极词汇?VADER代码异常排查
尝试从整段文本中识别积极、中立和消极词汇,使用VADER工具时,实际输出却是单个字母(比如Positive列表出现"a", "c", "y"等),无法正确提取段落中的词汇。
示例文本:
text1 = "Andromeda: is the 19th largest constellation in the sky.. It is located in the first quadrant of the northern hemisphere (NQ1).Andromeda has three stars brighter than magnitude 3.00 and three stars located within 10 parsecs (32.6 light years) of Earth. The brightest star in the constellation is Alpheratz. The nearest star is Ross 248 (spectral class M6V), also known as HH Andromedae, found at a distance of only 10.30 light years from Earth. The constellation is associated with the Andromedids meteor shower (also known as the Bielids), first documented on December 6, 1741 over Russia. The meteor shower has faded since discovery, but some activity is still observable in mid-November."
原错误代码:
nltk.download('vader_lexicon') from nltk.sentiment.vader import SentimentIntensityAnalyzer sid = SentimentIntensityAnalyzer() pos_word_list=[] neu_word_list=[] neg_word_list=[] for word in text1: if (sid.polarity_scores(word)['compound']) >= 0.5: pos_word_list.append(word) elif (sid.polarity_scores(word)['compound']) <= -0.5: neg_word_list.append(word) else: neu_word_list.append(word) print('Positive :',pos_word_list) print('Neutral :',neu_word_list) print('Negative :',neg_word_list)
原代码中for word in text1是遍历字符串的每个字符,而非拆分后的单词,导致VADER在单个字符上计算情感得分,最终输出都是单个字母。
需先将文本拆分为独立单词,再逐个分析每个单词的情感倾向。可以用nltk.word_tokenize()完成分词,同时可选择性清理标点、数字等无意义元素。
修正后的代码:
import nltk nltk.download('vader_lexicon') nltk.download('punkt') # 下载分词所需数据集 from nltk.sentiment.vader import SentimentIntensityAnalyzer from nltk.tokenize import word_tokenize sid = SentimentIntensityAnalyzer() text1 = "Andromeda: is the 19th largest constellation in the sky.. It is located in the first quadrant of the northern hemisphere (NQ1).Andromeda has three stars brighter than magnitude 3.00 and three stars located within 10 parsecs (32.6 light years) of Earth. The brightest star in the constellation is Alpheratz. The nearest star is Ross 248 (spectral class M6V), also known as HH Andromedae, found at a distance of only 10.30 light years from Earth. The constellation is associated with the Andromedids meteor shower (also known as the Bielids), first documented on December 6, 1741 over Russia. The meteor shower has faded since discovery, but some activity is still observable in mid-November." # 分词并过滤非字母内容(可根据需求调整) words = [word for word in word_tokenize(text1) if word.isalpha()] pos_word_list=[] neu_word_list=[] neg_word_list=[] for word in words: scores = sid.polarity_scores(word) compound_score = scores['compound'] if compound_score >= 0.5: pos_word_list.append(word) elif compound_score <= -0.5: neg_word_list.append(word) else: neu_word_list.append(word) print('Positive :', pos_word_list) print('Neutral :', neu_word_list) print('Negative :', neg_word_list)
word_tokenize()会将文本拆分为符合语言逻辑的单词,替代原有的字符遍历逻辑。- 过滤非字母内容是为了避免标点、数字这类无情感倾向的元素干扰结果,可根据实际需求决定是否保留。
- VADER的
compound得分范围为-1(极消极)到1(极积极),原代码的阈值(≥0.5和≤-0.5)可根据分析场景灵活调整。
内容的提问来源于stack exchange,提问作者Emily Sims

