You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用NLP与Python提取文本中的极端情感词汇?

用Python提取极端情感词汇的可行方案

下面提供几种实用方法帮你筛选出great、awesome这类极端情感词,排除good、like这类普通情感词:

方法1:基于VADER情感强度词典筛选

VADER是专门针对社交媒体文本的情感分析工具,能给词汇输出精准的情感强度评分,适合用来区分极端和普通情感词。

import nltk
from nltk.sentiment import SentimentIntensityAnalyzer

# 首次运行需下载VADER词典
nltk.download('vader_lexicon')

sia = SentimentIntensityAnalyzer()

# 待检测词汇列表
target_words = ["great", "awesome", "amazing", "interesting", "intriguing", "good", "bad", "like", "dislike"]

# 设定极端情感强度阈值(可根据需求调整)
EXTREME_THRESHOLD = 0.5

extreme_words = []
for word in target_words:
    score_dict = sia.polarity_scores(word)
    # 取正向/负向强度的最大值
    max_intensity = max(score_dict['pos'], score_dict['neg'])
    # 同时排除你指定的普通情感词
    if max_intensity >= EXTREME_THRESHOLD and word not in ["good", "bad", "like", "dislike"]:
        extreme_words.append(word)

print("筛选出的极端情感词汇:", extreme_words)

方法2:自定义极端情感词库匹配

如果你已经明确知道哪些属于极端情感词,直接维护一个词库做精准匹配是最可靠的方式:

# 自定义极端情感词库
extreme_positive = {"great", "awesome", "amazing", "interesting", "intriguing", "incredible", "fantastic"}
extreme_negative = {"terrible", "horrible", "awful", "devastating", "appalling"}

# 待检测词汇
input_words = ["great", "awesome", "good", "bad", "like", "dislike", "incredible"]

# 筛选匹配结果
extreme_words = [word for word in input_words if word in extreme_positive or word in extreme_negative]
print("筛选出的极端情感词汇:", extreme_words)

方法3:基于词向量扩展极端词库

如果需要自动扩展极端情感词范围,可以用预训练词向量(如Word2Vec)找到和已知极端词语义相近的词汇:

from gensim.models import KeyedVectors

# 加载预训练Word2Vec模型(需提前下载GoogleNews-vectors-negative300.bin)
model = KeyedVectors.load_word2vec_format('GoogleNews-vectors-negative300.bin', binary=True)

# 核心极端情感词
core_extreme = ["awesome", "amazing", "incredible"]
# 需要排除的普通情感词
common_exclude = {"good", "bad", "like", "dislike"}

# 语义相似度阈值
SIMILARITY_THRESHOLD = 0.6

extreme_words_set = set(core_extreme)
for word in core_extreme:
    # 取Top20语义相似词
    similar_pairs = model.most_similar(word, topn=20)
    for similar_word, score in similar_pairs:
        if score >= SIMILARITY_THRESHOLD and similar_word not in common_exclude:
            extreme_words_set.add(similar_word)

print("扩展后的极端情感词汇:", list(extreme_words_set))

内容的提问来源于stack exchange,提问作者vaibhav jain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 04:25:17