预训练DistilBERT文本分类多GPU并行实现报错求助
多GPU并行运行DistilBERT文本分类任务解决方案
错误原因分析
使用torch.nn.DataParallel包装模型后直接传给transformers的pipeline会报错,因为DataParallel实例不具备原模型的config属性,而pipeline依赖该属性完成标签映射等操作。
方案一:基于DataParallel的手动推理(数据并行,适合加速批量任务)
这种方式避开pipeline,手动处理tokenization和推理逻辑,直接使用DataParallel实现多GPU数据并行:
1. 修改模型加载函数
import torch from transformers import AutoModelForSequenceClassification, AutoTokenizer from collections import Counter import pandas as pd import numpy as np def load_nlp(directory_path, num_labels=2): num_gpus = torch.cuda.device_count() device = torch.device("cuda" if num_gpus > 0 else "cpu") tokenizer_ = AutoTokenizer.from_pretrained(directory_path) model_ = AutoModelForSequenceClassification.from_pretrained( directory_path, num_labels=num_labels) # 用DataParallel包装模型,分配到所有可用GPU if num_gpus > 1: model_ = torch.nn.DataParallel(model_, device_ids=list(range(num_gpus))) model_ = model_.to(device) return tokenizer_, model_
2. 重写分类推理函数
def label_cls(posts, tokenizer, model) -> str: # 对批量文本进行tokenize inputs = tokenizer( posts, padding=True, truncation=True, max_length=512, return_tensors="pt" ) # 将输入张量移到模型所在设备 inputs = {k: v.to(model.device) for k, v in inputs.items()} # 关闭梯度计算,加速推理 with torch.no_grad(): outputs = model(**inputs) # 获取预测标签ID,并映射为标签名 predictions = torch.argmax(outputs.logits, dim=1).cpu().numpy() # 从原模型(DataParallel包装后需通过module访问)获取id到标签的映射 if isinstance(model, torch.nn.DataParallel): id2label = model.module.config.id2label else: id2label = model.config.id2label labels = [id2label[pred] for pred in predictions] # 统计最常见标签 most_common_label = Counter(labels).most_common(1)[0][0] return most_common_label
3. 执行分类任务
saved_model_local_path = "model_path" tokenizer, model = load_nlp(saved_model_local_path, num_labels=2) # 对每个用户的帖子集合进行分类 user_agg["user_label"] = user_agg.user_posts_cat.apply( lambda posts: label_cls(posts, tokenizer, model) ) # 转换为机器人/人类标签 user_agg["user_label"] = np.where(user_agg.user_label == "LABEL_1", "bot", "human")
方案二:基于Accelerate的自动多GPU分配(模型并行,适合大模型)
如果你的模型较大,单GPU无法容纳,可以使用transformers的accelerate库实现自动模型层分配:
1. 安装依赖
pip install accelerate
2. 修改模型加载与推理代码
import torch from transformers import AutoModelForSequenceClassification, AutoTokenizer, pipeline from collections import Counter import pandas as pd import numpy as np def load_nlp(directory_path, num_labels=2): tokenizer_ = AutoTokenizer.from_pretrained(directory_path) # 使用device_map="auto"自动将模型层分配到多GPU model_ = AutoModelForSequenceClassification.from_pretrained( directory_path, num_labels=num_labels, device_map="auto" ) pipe = pipeline("text-classification", model=model_, tokenizer=tokenizer_) return pipe saved_model_local_path = "model_path" pipe = load_nlp(saved_model_local_path, num_labels=2) def label_cls(posts, pipe=pipe) -> str: labels = pipe(posts, padding=True, truncation=True, max_length=512) most_common_label = Counter([item["label"] for item in labels]).most_common(1)[0][0] return most_common_label user_agg["user_label"] = [label_cls(posts) for posts in user_agg.user_posts_cat] user_agg["user_label"] = np.where(user_agg.user_label == "LABEL_1", "bot", "human")
说明
- 方案一为数据并行:将同一份模型复制到所有GPU,每个GPU处理一部分批量数据,适合小模型加速批量推理任务。
- 方案二为模型并行:将模型的不同层分配到不同GPU,适合单个GPU无法容纳的大模型。
内容的提问来源于stack exchange,提问作者Kevin Li
相关产品推荐
相关产品推荐

