You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于LLM问答模型的语句情感分类需求及代码实现尝试

情感分类方案验证与优化指导

现有实现的核心问题

  1. 模型选型错误:你使用的是基于SQuAD数据集微调的BERT问答模型,这类模型的核心能力是从给定上下文中提取问题的答案,而非进行情感分类。它不会输出"积极/消极/中性"这类标签,反而会从输入句子里抓取相关词汇(比如"love""terrible"),完全不符合需求。
  2. 输入逻辑错误:问答任务需要明确的"问题+上下文"结构,但你的prompt将问题和待分类句子混在一起,上下文仅传入原句,模型无法理解要输出指定分类标签的要求。
  3. 冗余代码:导入了大量未使用的库(如Roberta相关组件、重复的AutoTokenizer),增加了不必要的资源占用。

优化方案

方案一:使用专门的情感分类预训练模型(推荐)

直接选用Hugging Face生态中针对情感分类任务微调的模型,这类模型能直接输出标准化的情感标签和置信度,准确率和效率远高于用问答模型适配。

示例代码:

import pandas as pd
from transformers import pipeline

# 初始化情感分类管道,采用轻量且高效的预训练模型
sentiment_analyzer = pipeline(
    "sentiment-analysis",
    model="distilbert-base-uncased-finetuned-sst-2-english"
)

# 构建数据集(保留你的原代码)
positive_sentences = [
    "I love this product!",
    "The weather is beautiful today.",
    "The team did an excellent job.",
    "She is a very talented musician."
]

negative_sentences = [
    "I am not satisfied with the service.",
    "The food was terrible at that restaurant.",
    "The movie was a complete disappointment.",
    "He made a lot of mistakes in the project."
]

sentences = positive_sentences + negative_sentences
df = pd.DataFrame({'snippet': sentences})

# 批量处理情感分类,效率远高于逐行迭代
results = sentiment_analyzer(df['snippet'].tolist())

# 将结果合并到DataFrame中
df['sentiment'] = [res['label'].lower() for res in results]
df['confidence'] = [res['score'] for res in results]

print(df)

方案二:适配问答模型实现情感分类(不推荐)

如果必须使用现有问答模型,需要重构输入逻辑,明确告知模型可选的情感标签范围:

示例代码:

import pandas as pd
from transformers import pipeline, BertTokenizer, BertForQuestionAnswering

# 加载问答模型
tokenizer = BertTokenizer.from_pretrained('bert-large-uncased-whole-word-masking-finetuned-squad')
model = BertForQuestionAnswering.from_pretrained('bert-large-uncased-whole-word-masking-finetuned-squad')
question_answerer = pipeline("question-answering", model=model, tokenizer=tokenizer)

# 构建数据集
positive_sentences = [
    "I love this product!",
    "The weather is beautiful today.",
    "The team did an excellent job.",
    "She is a very talented musician."
]

negative_sentences = [
    "I am not satisfied with the service.",
    "The food was terrible at that restaurant.",
    "The movie was a complete disappointment.",
    "He made a lot of mistakes in the project."
]

sentences = positive_sentences + negative_sentences
df = pd.DataFrame({'snippet': sentences})

# 重构prompt与上下文,明确给出可选标签
for index, row in df.iterrows():
    sentence = row["snippet"]
    context = f"Sentence: {sentence}. The possible sentiment labels are positive, negative, neutral."
    question = "What is the sentiment of the sentence?"
    
    result = question_answerer(question=question, context=context)
    
    if "answer" in result:
        print(f"Article {index+1} sentiment: {result['answer']}")
    else:
        print(f"Article {index+1} sentiment not found.")

额外优化建议

  • 清理冗余代码:删除未使用的库导入,减少内存占用。
  • 支持中性分类:若需要区分中性情感,可选用支持三分类的模型(如cardiffnlp/twitter-roberta-base-sentiment),该模型输出LABEL_0(消极)、LABEL_1(中性)、LABEL_2(积极),可自行映射为中文标签。
  • 轻量化部署:在资源有限的环境中,优先选用DistilBERT、MobileBERT等轻量化模型,兼顾速度与准确率。
  • 批量处理:避免使用iterrows()逐行处理,改用批量输入提升效率。

内容的提问来源于stack exchange,提问作者Economist_Ayahuasca

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 14:03:29