如何让arphanghoshal/EmoRoBERTa模型实现秒级响应?
优化arphanghoshal/EmoRoBERTa模型运行速度的方案
原问题场景
运行arpanghoshal/EmoRoBERTa模型时响应极慢,原调用代码如下:
from transformers import RobertaTokenizerFast, TFRobertaForSequenceClassification, pipeline tokenizer = RobertaTokenizerFast.from_pretrained("arpanghoshal/EmoRoBERTa") model = TFRobertaForSequenceClassification.from_pretrained("arpanghoshal/EmoRoBERTa") emotxt = pipeline('sentiment-analysis', model='arpanghoshal/EmoRoBERTa') emotions = emotxt(query) emotions
核心优化方案
以下是几个能快速将响应时间压缩到数秒内的优化手段:
1. 避免重复加载模型,指定硬件加速设备
原代码重复加载模型实例,且未明确指定GPU设备,导致资源浪费和CPU推理低效。修正后的基础代码:
from transformers import RobertaTokenizerFast, RobertaForSequenceClassification, pipeline # 优先使用PyTorch版本模型(比TensorFlow版本推理更快) model = RobertaForSequenceClassification.from_pretrained("arpanghoshal/EmoRoBERTa") tokenizer = RobertaTokenizerFast.from_pretrained("arpanghoshal/EmoRoBERTa") # 指定device=0使用GPU,无GPU则用device=-1 emotxt = pipeline('sentiment-analysis', model=model, tokenizer=tokenizer, device=0) emotions = emotxt(query) print(emotions)
2. 启用模型量化
通过量化将模型权重从FP32转为INT8/FP16,大幅减少显存占用并提升推理速度,无需修改模型结构:
from transformers import RobertaTokenizerFast, RobertaForSequenceClassification, pipeline from transformers import BitsAndBytesConfig # 配置4-bit量化 bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16 ) model = RobertaForSequenceClassification.from_pretrained( "arpanghoshal/EmoRoBERTa", quantization_config=bnb_config, device_map="auto" ) tokenizer = RobertaTokenizerFast.from_pretrained("arpanghoshal/EmoRoBERTa") emotxt = pipeline('sentiment-analysis', model=model, tokenizer=tokenizer) emotions = emotxt(query) print(emotions)
3. 使用ONNX Runtime加速
将模型导出为ONNX格式,利用ONNX Runtime的优化推理引擎进一步提速:
from transformers import RobertaTokenizerFast, RobertaForSequenceClassification from transformers.onnx import export import onnxruntime as ort import torch # 导出模型为ONNX model = RobertaForSequenceClassification.from_pretrained("arpanghoshal/EmoRoBERTa") tokenizer = RobertaTokenizerFast.from_pretrained("arpanghoshal/EmoRoBERTa") onnx_path = "emoroberta.onnx" export(tokenizer, model, opset=17, output=onnx_path) # 用ONNX Runtime加载并推理 session = ort.InferenceSession(onnx_path) inputs = tokenizer(query, return_tensors="np") # 执行推理 outputs = session.run(None, dict(inputs)) predicted_class_idx = outputs[0].argmax(axis=1)[0] emotion = model.config.id2label[predicted_class_idx] print({"label": emotion, "score": float(outputs[0][0][predicted_class_idx])})
4. 批量处理文本(多query场景)
如果需要处理多个文本,批量传入pipeline而非逐个调用,减少重复的tokenizer和模型调用开销:
# 示例:批量处理多个query queries = ["我今天很开心", "这件事让我很沮丧", "这个结果太意外了"] emotions = emotxt(queries) print(emotions)
内容的提问来源于stack exchange,提问作者Abdullah
相关产品推荐
相关产品推荐

