You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于EmoRoBERTa的情感词提取代码性能优化方案咨询

问题描述

我用以下Python代码,借助EmoRoBERTa的emotion_pipeline提取文本中的情感相关词汇:

temp_list = processed_text.split()
print(temp_list)

emotion_lists = {
    'admiration': [], 'amusement': [], 'anger': [], 'annoyance': [],
    'approval': [], 'caring': [], 'confusion': [], 'curiosity': [],
    'desire': [], 'disappointment': [], 'disapproval': [], 'disgust': [],
    'embarrassment': [], 'excitement': [], 'fear': [], 'gratitude': [],
    'grief': [], 'joy': [], 'love': [], 'nervousness': [], 'optimism': [],
    'pride': [], 'realization': [], 'relief': [], 'remorse': [], 'sadness': [],
    'surprise': []
}

for x in temp_list:
    results = emotion_pipeline(x) # 基于EmoRoBERTa的情感预测Pipeline
    if results != 'neutral':
        for result in results:
            top_label = max(result, key=lambda dictionary: dictionary['score'])
            label = top_label["label"]
            if label in emotion_lists:
                emotion_lists[label].append(x)
            
print(emotion_lists)

但处理长段落时代码执行速度过慢,请问如何进一步优化?

已做优化

我已移除冗余代码,添加条件语句跳过中性词处理,并从emotion_list中移除中性词列表,将执行时间从1分45秒缩短至1分22秒。

基准测试文本

今天,我想花点时间反思生命之美,分享一些最近萦绕在我心头的想法。

生命是由经历、情感与联结编织而成的复杂织锦。这是一段充满起起落落、曲折变化与意外惊喜的旅程。人们很容易陷入日常琐事,忽略身边的美好。但正是在这些反思的时刻,我们才能真正领略生命的壮丽。

想想那些给你带来快乐的简单小事:寒冷清晨的一杯热咖啡、爱人温柔的抚摸,或是日落时分令人惊叹的色彩。这些看似平凡的瞬间,正是让生命变得不凡的原因。培养感恩之心,我们就能在最微小的事物中找到美好与满足。

生命也关乎成长与学习。每一次经历,无论是愉快还是充满挑战,都为个人成长提供了机会。正是在逆境中,我们才能发现自己的力量与韧性。拥抱变化、走出舒适区,能带来深刻的转变与全新的视野。记住,重要的不是目的地,而是旅程本身。

此外,生命离不开联结与关系。我们与他人的互动塑造了我们的经历,也让我们有归属感。花点时间珍惜那些触动你生命的人,无论是家人、朋友,还是曾施以善意的陌生人。用心维系这些关系,因为它们能给予我们支持、理解,还有共同欢笑与落泪的时刻。

尽管生命充满不确定性与挑战,但保持希望与乐观至关重要。拥抱梦想与抱负的力量,它们能点燃我们的热情,推动我们前进。庆祝每一个成就,无论大小,把挫折当作迈向更大成功的垫脚石。

总之,生命是一份珍贵的礼物,如何充分利用它取决于我们自己。让我们以感恩、好奇与开放的心态迎接每一天。拥抱身边的美好,寻求个人成长,维系人际关系,坚守希望。记住,生命的真谛不在于长度,而在于经历的深度。

emotion_pipeline执行时间

emotion_pipeline执行时间

emotion_pipeline单词输出示例(temp_list最后一个词)

[[{'label': 'admiration', 'score': 0.0011110481573268771}, {'label': 'amusement', 'score': 0.0023590335622429848}, {'label': 'anger', 'score': 8.838719077175483e-05}, {'label': 'annoyance', 'score': 0.00025082440697588027}, {'label': 'approval', 'score': 0.0007353303371928632}, {'label': 'caring', 'score': 4.328513750806451e-05}, {'label': 'confusion', 'score': 5.098163819639012e-05}, {'label': 'curiosity', 'score': 3.086668220930733e-05}, {'label': 'desire', 'score': 4.53599714091979e-05}, {'label': 'disappointment', 'score': 7.120604277588427e-05}, {'label': 'disapproval', 'score': 0.00014575349632650614}, {'label': 'disgust', 'score': 0.00015191144484560937}, {'label': 'embarrassment', 'score': 2.4345894416910596e-05}, {'label': 'excitement', 'score': 5.621055970550515e-05}, {'label': 'fear', 'score': 3.3071850339183584e-05}, {'label': 'gratitude', 'score': 1.76055655174423e-05}, {'label': 'grief', 'score': 1.189570775750326e-05}, {'label': 'joy', 'score': 8.78259088494815e-05}, {'label': 'love', 'score': 4.430738772498444e-05}, {'label': 'nervousness', 'score': 1.218793750012992e-05}, {'label': 'optimism', 'score': 5.256943040876649e-05}, {'label': 'pride', 'score': 2.5214694687747397e-05}, {'label': 'realization', 'score': 9.513740224065259e-05}, {'label': 'relief', 'score': 6.58453109281254e-06}, {'label': 'remorse', 'score': 4.798649752046913e-05}, {'label': 'sadness', 'score': 0.00011766400712076575}, {'label': 'surprise', 'score': 2.1648049369105138e-05}, {'label': 'neutral', 'score': 0.9942617416381836}]]

优化方案

1. 批量处理单词,减少模型调用开销

当前代码逐个调用emotion_pipeline,每次调用都存在模型加载、推理的固定开销。transformers的Pipeline支持直接传入单词列表批量处理,能大幅降低调用次数,提升执行效率。

修改后的代码:

temp_list = processed_text.split()
# 先去重处理,避免重复计算;若需保留原文本中所有单词出现的次数,后续再映射回去
unique_words = list(set(temp_list))

emotion_lists = {
    'admiration': [], 'amusement': [], 'anger': [], 'annoyance': [],
    'approval': [], 'caring': [], 'confusion': [], 'curiosity': [],
    'desire': [], 'disappointment': [], 'disapproval': [], 'disgust': [],
    'embarrassment': [], 'excitement': [], 'fear': [], 'gratitude': [],
    'grief': [], 'joy': [], 'love': [], 'nervousness': [], 'optimism': [],
    'pride': [], 'realization': [], 'relief': [], 'remorse': [], 'sadness': [],
    'surprise': []
}

# 批量传入所有单词,设置return_all_scores=False只返回最高分标签,减少数据处理量
results = emotion_pipeline(unique_words, return_all_scores=False)

# 遍历结果,将单词映射到对应情感列表(保留原文本中的出现次数)
for word, result in zip(unique_words, results):
    label = result['label']
    if label != 'neutral' and label in emotion_lists:
        # 按原文本中单词出现的次数添加
        emotion_lists[label].extend([word] * temp_list.count(word))

print(emotion_lists)

2. 缓存重复单词的推理结果

若不想去重(比如需要保留单词在原文本中的顺序),可以用字典缓存已处理过的单词结果,避免重复计算:

temp_list = processed_text.split()
emotion_lists = {
    # 同之前的情感字典定义
}
result_cache = {}

for word in temp_list:
    if word in result_cache:
        label = result_cache[word]
    else:
        # 只返回最高分标签,减少数据处理
        result = emotion_pipeline(word, return_all_scores=False)[0]
        label = result['label']
        result_cache[word] = label
    
    if label != 'neutral' and label in emotion_lists:
        emotion_lists[label].append(word)

print(emotion_lists)

3. 提前过滤无意义停用词

文本中的停用词(如"the"、"is"、"and")大多属于中性,提前过滤可减少需要处理的单词数量。可以用nltk的停用词库:

import nltk
nltk.download('stopwords')
from nltk.corpus import stopwords

stop_words = set(stopwords.words('english'))
# 过滤停用词,同时保留大小写形式(若需统一小写可添加word.lower())
temp_list = [word for word in processed_text.split() if word.lower() not in stop_words]

# 后续处理同批量或缓存方案

4. 启用GPU加速

如果有GPU资源,初始化Pipeline时指定设备,让模型在GPU上推理,速度可提升数倍:

from transformers import pipeline

# device=0表示使用第一个GPU,无GPU时会自动切换为CPU
emotion_pipeline = pipeline("text-classification", model="SamLowe/roberta-base-go_emotions", device=0)

5. 使用量化模型加速推理

借助bitsandbytes库对模型做4位/8位量化,减少显存占用的同时提升推理速度:

from transformers import pipeline, AutoModelForSequenceClassification, AutoTokenizer
import torch

model = AutoModelForSequenceClassification.from_pretrained(
    "SamLowe/roberta-base-go_emotions",
    load_in_4bit=True,
    device_map="auto",
    torch_dtype=torch.bfloat16
)
tokenizer = AutoTokenizer.from_pretrained("SamLowe/roberta-base-go_emotions")
emotion_pipeline = pipeline("text-classification", model=model, tokenizer=tokenizer)

内容的提问来源于stack exchange,提问作者abdz_128

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 15:08:09