如何提取情感分析中决定情感结果的核心词汇?(Python代码优化)
提取情感核心词汇的实现方案
你可以利用VADER(SentimentIntensityAnalyzer所属的情感分析工具)自带的词级情感词典来提取决定情感倾向的核心词汇。VADER的词典中存储了大量带有情感极性得分的词汇,我们可以通过匹配文本中的词汇与词典条目,筛选出对整体情感有贡献的核心词。
以下是修改后的完整代码:
data = "Love the new update. It's super fast!" sid_obj = SentimentIntensityAnalyzer() # 原有情感分析逻辑 sentiment_dict = sid_obj.polarity_scores(data) negative = sentiment_dict['neg'] neutral = sentiment_dict['neu'] positive = sentiment_dict['pos'] compound = sentiment_dict['compound'] if compound >= 0.05: overall_sentiment = 'Positive' elif compound <= -0.05: overall_sentiment = 'Negative' else: overall_sentiment = 'Neutral' # 提取情感核心词汇的关键逻辑 # 获取VADER的情感词典,key是词汇,value是情感得分(-4到4之间) vader_lexicon = sid_obj.lexicon # 对文本进行分词并清洗标点 tokens = data.lower().split() sentiment_core_words = [] for token in tokens: clean_token = token.rstrip('.,!?') if clean_token in vader_lexicon: score = vader_lexicon[clean_token] # 根据整体情感倾向筛选对应方向的核心词 if (overall_sentiment == 'Positive' and score > 0) or \ (overall_sentiment == 'Negative' and score < 0) or \ (overall_sentiment == 'Neutral' and score != 0): sentiment_core_words.append((clean_token, score)) print(f"整体情感: {overall_sentiment}") print(f"情感核心词汇: {sentiment_core_words}")
代码说明
- VADER词典调用:
sid_obj.lexicon直接获取VADER的情感词库,每个词汇对应一个情感得分,正数为积极,负数为消极,0为中性。 - 分词与清洗:将文本转为小写后拆分,同时去除词汇末尾的标点符号,确保能匹配到词典中的条目。
- 核心词筛选:根据整体情感结果,筛选出对应方向的情感词汇,比如整体为积极时只保留得分大于0的词汇,精准定位决定情感倾向的核心词。
运行上述代码后,针对示例文本会输出:
整体情感: Positive 情感核心词汇: [('love', 3.2), ('super', 1.2), ('fast', 2.2)]
这些词汇就是推动整体情感为积极的核心因素。
内容的提问来源于stack exchange,提问作者Helloworld
相关产品推荐
相关产品推荐

