基于自定义情感词典实现文本情感分析的方法咨询
情感分析功能实现方案
你可以按以下逻辑完成后续功能开发:
- 兼容大小写匹配:测试样例中存在首字母大写的情感词(如
Good、Excellent),需要将分词统一转小写后再和词表匹配,避免匹配遗漏。 - 定义情感词表:建议用集合类型存储正负向词,查询匹配效率高于列表。
- 编写情感判断函数:遍历预处理后的分词列表,分别统计正负向词的出现次数,比较计数结果后返回对应情感标识。
完整可运行代码如下:
import string # 已有的文本预处理函数 def splitandremovepunc(s): t = s.maketrans("", "", string.punctuation) return s.translate(t).split() # 定义正负向情感词表 positive_words = {"good", "awesome", "excellent", "great"} negative_words = {"bad", "broke", "terrible", "poor"} # 情感判断函数 def get_sentiment(processed_words, pos_vocab, neg_vocab): pos_count = 0 neg_count = 0 for word in processed_words: lower_word = word.lower() if lower_word in pos_vocab: pos_count += 1 elif lower_word in neg_vocab: neg_count += 1 # 正向词数量更多则判定为正向,其余情况判定为负向,平局判定规则可自行调整 return "positive comment" if pos_count > neg_count else "negative comment" # 测试评论 lst = "Good for the price, but poor Bluetooth connections." lst2 = "Excellent product. Awesome quality and good customer service." lst3 = "The quality is terrible. I would not buy this product again." # 批量处理并输出结果 for comment in [lst, lst2, lst3]: processed = splitandremovepunc(comment) print(get_sentiment(processed, positive_words, negative_words))
运行后输出结果:
negative comment positive comment negative comment
如果需要调整正负向词数量相等时的判定规则,修改get_sentiment函数里的判断条件即可。
内容的提问来源于stack exchange,提问作者Asbjørn Vattøy
相关产品推荐
相关产品推荐

