如何统计带有正面/负面/中性情感评分的句子数量
统计不同情感倾向句子数量的解决方案
我来帮你搞定这个情感句子统计的问题~看你现在的代码思路方向是对的,但有几个关键细节需要调整,才能正确统计单个句子的情感并完成计数:
现有代码的核心问题
- 你把
token_dictionary初始化放在了句子循环外面,会导致所有句子的token情感值被累加在一起,根本没法区分单个句子的情感倾向 - 用列表来存计数(比如
neg_count = [])是错误的,计数应该用整数变量来累加 - 代码没完成单个句子的情感得分计算和最终的倾向判断逻辑
修正后的完整代码
from nltk.tokenize import sent_tokenize, word_tokenize def sent_dictionary__sentence_(text): # 初始化计数变量,用整数替代列表才是正确的累加方式 neg_count = 0 pos_count = 0 neu_count = 0 lexicon_keys = lexicon_dictionary.keys() sentences = sent_tokenize(text) for sentence in sentences: # 每个句子单独初始化token字典,避免跨句子的情感值混淆 token_dictionary = {} # 按流程处理句子:分词→去标点→词形还原 tokens = apply_lemmatization(remove_punctuation(word_tokenize(str(sentence)))) for token in tokens: if token in lexicon_keys: # 累加当前句子中该token的情感值 token_dictionary[token] = token_dictionary.get(token, 0) + lexicon_dictionary[token] # 计算当前句子的总情感得分 sentiment_score = sum(token_dictionary.values()) # 根据得分判断情感倾向并完成计数 if sentiment_score > 0: pos_count += 1 elif sentiment_score < 0: neg_count += 1 else: neu_count += 1 # 返回结构化的计数结果 return {"positive_sentences": pos_count, "negative_sentences": neg_count, "neutral_sentences": neu_count}
关键调整说明
- 单句独立计算:把
token_dictionary的初始化移到句子循环内部,确保每个句子的情感计算都是独立的,不会和其他句子的结果混在一起 - 正确计数逻辑:用整数变量替代列表,每次判断完句子情感后直接累加对应计数
- 情感判断规则:通过单个句子的情感得分总和来判定倾向——得分大于0为正面,小于0为负面,等于0为中性(你可以根据自己的需求调整阈值,比如设置一个微小区间来过滤得分接近0的模糊情况)
假设你已经正确定义了lexicon_dictionary(单词到情感值的映射字典)、apply_lemmatization(词形还原函数)和remove_punctuation(去标点函数),这个函数就能准确返回各情感倾向的句子数量了。
内容的提问来源于stack exchange,提问作者Shelina
相关产品推荐
相关产品推荐

