PySpark情感分析误判问题:如何优化代码提升准确率?
问题分析与修正方案
你的核心问题出在停用词过滤逻辑错误和模型适配性不足,导致带转折的正面评论被误判:
- 停用词过滤时移除了
but这类关键逻辑连接词,模型无法识别句子的转折关系,只看到前半段负面描述就给出负面判断; - 所用的
distilbert-base-uncased-finetuned-sst-2-english是针对简单电影评论训练的,对真实客户评论里的拼写错误、复杂转折场景鲁棒性差。
具体修正步骤
1. 调整停用词过滤逻辑,保留关键连接词
把but、however这类转折词从停用词集合中排除,这些词是判断句子情感倾向的核心逻辑标识,不能过滤。
2. 优化预处理,处理拼写错误(可选但有效)
针对像grate(应为great)这类拼写错误,加入简单的拼写校正逻辑,减少模型的识别偏差。
3. 更换适配真实评论场景的模型
改用在亚马逊评论数据集上微调的模型,这类模型对客户服务类评论的情感判断更准确,比如distilbert-base-uncased-finetuned-amazon-review-classification。
修正后的代码
from pyspark.sql.functions import udf, col from pyspark.sql.types import StringType from transformers import pipeline from nltk.corpus import stopwords from nltk.tokenize import word_tokenize from spellchecker import SpellChecker # 需先安装:pip install pyspellchecker # 初始化拼写检查器 spell = SpellChecker() # 定义过滤停用词函数,保留关键逻辑连接词 def filter_stopwords(sentence): stop_words = set(stopwords.words('english')) # 移除转折类连接词,保留逻辑关系 keep_words = {"but", "however", "yet", "though", "although"} stop_words = stop_words - keep_words word_tokens = word_tokenize(sentence) filtered_sentence = [w for w in word_tokens if not w in stop_words] return " ".join(filtered_sentence) # 拼写校正函数 def correct_spelling(text): words = text.split() corrected_words = [] for word in words: # 跳过标点和专有名词 if word.isalpha(): corrected_word = spell.correction(word) corrected_words.append(corrected_word if corrected_word else word) else: corrected_words.append(word) return " ".join(corrected_words) # 初始化针对亚马逊评论优化的情感分析模型 sentiment_pipeline = pipeline( "sentiment-analysis", model="distilbert-base-uncased-finetuned-amazon-review-classification" ) # 定义情感分析函数,整合预处理步骤 def get_sentiment(text): # 先校正拼写 corrected_text = correct_spelling(text) # 过滤停用词(保留关键连接词) filtered_text = filter_stopwords(corrected_text) result = sentiment_pipeline(filtered_text)[0] # 将模型输出转换为统一的positive/negative return "positive" if result['label'] == "POSITIVE" else "negative" # 注册UDF sentiment_udf = udf(get_sentiment, StringType()) # 使用示例 # df = df.withColumn("sentiment", sentiment_udf(col("text_column")))
测试用例验证说明
- 对于句子1/3:保留
but后,模型能识别前半段负面是铺垫,后半段正面是核心判断,会输出positive; - 对于句子2:拼写校正后
grate变为great,模型能正确识别正面情感。
内容的提问来源于stack exchange,提问作者sande
相关产品推荐
相关产品推荐

