如何用Python基于分词与情感词典实现情感分析训练数据二分类标注
基于情感词典的情感二分类标注Python实现
前置依赖安装
根据你处理的文本语言安装对应分词工具:
- 中文场景安装
jieba分词:pip install jieba - 英文场景安装
nltk:pip install nltk,首次运行前执行nltk.download('punkt')下载分词数据包
完整实现代码
以下是中文场景的可运行示例,处理英文仅需替换分词逻辑:
import jieba # 加载本地情感词典,替换为你自己的词典路径即可 def load_sentiment_dict(pos_dict_path, neg_dict_path): # 读取正向词表,用set存储可实现O(1)速度查询 with open(pos_dict_path, 'r', encoding='utf-8') as f: pos_words = set([line.strip() for line in f if line.strip()]) # 读取负向词表 with open(neg_dict_path, 'r', encoding='utf-8') as f: neg_words = set([line.strip() for line in f if line.strip()]) return pos_words, neg_words # 单条文本情感标注函数 def get_sentiment_label(text, pos_words, neg_words): # 分词并过滤空值 tokens = [token for token in jieba.lcut(text) if token.strip()] # 统计正负向词出现次数 pos_count = sum(1 for token in tokens if token in pos_words) neg_count = sum(1 for token in tokens if token in neg_words) # 比较次数输出标签,次数相等的规则可自行调整 if pos_count > neg_count: return "positive" elif neg_count > pos_count: return "negative" else: # 必须二分类场景可自行修改返回值,或者过滤这类中性数据 return "neutral" # 批量标注示例 if __name__ == "__main__": # 替换为你的情感词典路径 pos_words, neg_words = load_sentiment_dict("positive_dict.txt", "negative_dict.txt") # 替换为你的待标注文档列表 doc_list = [ "这款产品做工精细,用起来体验特别好", "什么垃圾质量,用了两天就坏了,售后也没人理", "今天气温25度,外面风很大" ] # 批量打标 labeled_result = [] for doc in doc_list: label = get_sentiment_label(doc, pos_words, neg_words) labeled_result.append((doc, label)) # 打印结果 for item in labeled_result: print(f"文本:{item[0]} | 标签:{item[1]}")
英文场景适配
处理英文时仅需要替换get_sentiment_label函数中的分词逻辑即可:
from nltk.tokenize import word_tokenize # 把原分词行替换为下面的代码,统一转小写匹配词典 tokens = [token.lower() for token in word_tokenize(text) if token.strip()]
优化提示
- 如果你的情感词典自带权重值,把统计次数的逻辑改为累加对应权重即可
- 正负词数量相等的场景,可额外加入否定词匹配逻辑优化标注准确率
内容的提问来源于stack exchange,提问作者Arkan
相关产品推荐
相关产品推荐

