You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从段落中提取积极/中立/消极词汇?VADER代码异常排查

问题描述

尝试从整段文本中识别积极、中立和消极词汇,使用VADER工具时,实际输出却是单个字母(比如Positive列表出现"a", "c", "y"等),无法正确提取段落中的词汇。

示例文本:

text1 = "Andromeda: is the 19th largest constellation in the sky.. It is located in the first quadrant of the northern hemisphere (NQ1).Andromeda has three stars brighter than magnitude 3.00 and three stars located within 10 parsecs (32.6 light years) of Earth. The brightest star in the constellation is Alpheratz. The nearest star is Ross 248 (spectral class M6V), also known as HH Andromedae, found at a distance of only 10.30 light years from Earth. The constellation is associated with the Andromedids meteor shower (also known as the Bielids), first documented on December 6, 1741 over Russia. The meteor shower has faded since discovery, but some activity is still observable in mid-November."

原错误代码:

nltk.download('vader_lexicon')
from nltk.sentiment.vader import SentimentIntensityAnalyzer

sid = SentimentIntensityAnalyzer()

pos_word_list=[]
neu_word_list=[]
neg_word_list=[]
for word in text1:
    if (sid.polarity_scores(word)['compound']) >= 0.5:
        pos_word_list.append(word)
    elif (sid.polarity_scores(word)['compound']) <= -0.5:
        neg_word_list.append(word) else: neu_word_list.append(word)

print('Positive :',pos_word_list) print('Neutral :',neu_word_list) print('Negative :',neg_word_list)
错误原因

原代码中for word in text1是遍历字符串的每个字符,而非拆分后的单词,导致VADER在单个字符上计算情感得分,最终输出都是单个字母。

解决方案

需先将文本拆分为独立单词,再逐个分析每个单词的情感倾向。可以用nltk.word_tokenize()完成分词,同时可选择性清理标点、数字等无意义元素。

修正后的代码:

import nltk
nltk.download('vader_lexicon')
nltk.download('punkt')  # 下载分词所需数据集
from nltk.sentiment.vader import SentimentIntensityAnalyzer
from nltk.tokenize import word_tokenize

sid = SentimentIntensityAnalyzer()

text1 = "Andromeda: is the 19th largest constellation in the sky.. It is located in the first quadrant of the northern hemisphere (NQ1).Andromeda has three stars brighter than magnitude 3.00 and three stars located within 10 parsecs (32.6 light years) of Earth. The brightest star in the constellation is Alpheratz. The nearest star is Ross 248 (spectral class M6V), also known as HH Andromedae, found at a distance of only 10.30 light years from Earth. The constellation is associated with the Andromedids meteor shower (also known as the Bielids), first documented on December 6, 1741 over Russia. The meteor shower has faded since discovery, but some activity is still observable in mid-November."

# 分词并过滤非字母内容(可根据需求调整)
words = [word for word in word_tokenize(text1) if word.isalpha()]

pos_word_list=[]
neu_word_list=[]
neg_word_list=[]

for word in words:
    scores = sid.polarity_scores(word)
    compound_score = scores['compound']
    if compound_score >= 0.5:
        pos_word_list.append(word)
    elif compound_score <= -0.5:
        neg_word_list.append(word)
    else:
        neu_word_list.append(word)

print('Positive :', pos_word_list)
print('Neutral :', neu_word_list)
print('Negative :', neg_word_list)
补充说明
  1. word_tokenize()会将文本拆分为符合语言逻辑的单词,替代原有的字符遍历逻辑。
  2. 过滤非字母内容是为了避免标点、数字这类无情感倾向的元素干扰结果,可根据实际需求决定是否保留。
  3. VADER的compound得分范围为-1(极消极)到1(极积极),原代码的阈值(≥0.5和≤-0.5)可根据分析场景灵活调整。

内容的提问来源于stack exchange,提问作者Emily Sims

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 19:25:12