You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark情感分析误判问题:如何优化代码提升准确率?

问题分析与修正方案

你的核心问题出在停用词过滤逻辑错误和模型适配性不足,导致带转折的正面评论被误判:

  1. 停用词过滤时移除了but这类关键逻辑连接词,模型无法识别句子的转折关系,只看到前半段负面描述就给出负面判断;
  2. 所用的distilbert-base-uncased-finetuned-sst-2-english是针对简单电影评论训练的,对真实客户评论里的拼写错误、复杂转折场景鲁棒性差。

具体修正步骤

1. 调整停用词过滤逻辑,保留关键连接词

把but、however这类转折词从停用词集合中排除,这些词是判断句子情感倾向的核心逻辑标识,不能过滤。

2. 优化预处理,处理拼写错误(可选但有效)

针对像grate(应为great)这类拼写错误,加入简单的拼写校正逻辑,减少模型的识别偏差。

3. 更换适配真实评论场景的模型

改用在亚马逊评论数据集上微调的模型,这类模型对客户服务类评论的情感判断更准确,比如distilbert-base-uncased-finetuned-amazon-review-classification。

修正后的代码

from pyspark.sql.functions import udf, col
from pyspark.sql.types import StringType
from transformers import pipeline
from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize
from spellchecker import SpellChecker  # 需先安装:pip install pyspellchecker

# 初始化拼写检查器
spell = SpellChecker()

# 定义过滤停用词函数,保留关键逻辑连接词
def filter_stopwords(sentence):
    stop_words = set(stopwords.words('english'))
    # 移除转折类连接词,保留逻辑关系
    keep_words = {"but", "however", "yet", "though", "although"}
    stop_words = stop_words - keep_words
    word_tokens = word_tokenize(sentence)
    filtered_sentence = [w for w in word_tokens if not w in stop_words]
    return " ".join(filtered_sentence)

# 拼写校正函数
def correct_spelling(text):
    words = text.split()
    corrected_words = []
    for word in words:
        # 跳过标点和专有名词
        if word.isalpha():
            corrected_word = spell.correction(word)
            corrected_words.append(corrected_word if corrected_word else word)
        else:
            corrected_words.append(word)
    return " ".join(corrected_words)

# 初始化针对亚马逊评论优化的情感分析模型
sentiment_pipeline = pipeline(
    "sentiment-analysis", 
    model="distilbert-base-uncased-finetuned-amazon-review-classification"
)

# 定义情感分析函数,整合预处理步骤
def get_sentiment(text):
    # 先校正拼写
    corrected_text = correct_spelling(text)
    # 过滤停用词(保留关键连接词)
    filtered_text = filter_stopwords(corrected_text)
    result = sentiment_pipeline(filtered_text)[0]
    # 将模型输出转换为统一的positive/negative
    return "positive" if result['label'] == "POSITIVE" else "negative"

# 注册UDF
sentiment_udf = udf(get_sentiment, StringType())

# 使用示例
# df = df.withColumn("sentiment", sentiment_udf(col("text_column")))

测试用例验证说明

  • 对于句子1/3:保留but后,模型能识别前半段负面是铺垫,后半段正面是核心判断,会输出positive;
  • 对于句子2:拼写校正后grate变为great,模型能正确识别正面情感。

内容的提问来源于stack exchange,提问作者sande

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 06:44:52