You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

预训练DistilBERT文本分类多GPU并行实现报错求助

多GPU并行运行DistilBERT文本分类任务解决方案

错误原因分析

使用torch.nn.DataParallel包装模型后直接传给transformers的pipeline会报错,因为DataParallel实例不具备原模型的config属性,而pipeline依赖该属性完成标签映射等操作。

方案一:基于DataParallel的手动推理(数据并行,适合加速批量任务)

这种方式避开pipeline,手动处理tokenization和推理逻辑,直接使用DataParallel实现多GPU数据并行:

1. 修改模型加载函数

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
from collections import Counter
import pandas as pd
import numpy as np

def load_nlp(directory_path, num_labels=2):
    num_gpus = torch.cuda.device_count()
    device = torch.device("cuda" if num_gpus > 0 else "cpu")

    tokenizer_ = AutoTokenizer.from_pretrained(directory_path)
    model_ = AutoModelForSequenceClassification.from_pretrained(
        directory_path, num_labels=num_labels)
    
    # 用DataParallel包装模型,分配到所有可用GPU
    if num_gpus > 1:
        model_ = torch.nn.DataParallel(model_, device_ids=list(range(num_gpus)))
    model_ = model_.to(device)

    return tokenizer_, model_

2. 重写分类推理函数

def label_cls(posts, tokenizer, model) -> str:
    # 对批量文本进行tokenize
    inputs = tokenizer(
        posts, 
        padding=True, 
        truncation=True, 
        max_length=512, 
        return_tensors="pt"
    )
    # 将输入张量移到模型所在设备
    inputs = {k: v.to(model.device) for k, v in inputs.items()}
    
    # 关闭梯度计算,加速推理
    with torch.no_grad():
        outputs = model(**inputs)
    
    # 获取预测标签ID,并映射为标签名
    predictions = torch.argmax(outputs.logits, dim=1).cpu().numpy()
    # 从原模型(DataParallel包装后需通过module访问)获取id到标签的映射
    if isinstance(model, torch.nn.DataParallel):
        id2label = model.module.config.id2label
    else:
        id2label = model.config.id2label
    
    labels = [id2label[pred] for pred in predictions]
    # 统计最常见标签
    most_common_label = Counter(labels).most_common(1)[0][0]
    return most_common_label

3. 执行分类任务

saved_model_local_path = "model_path"
tokenizer, model = load_nlp(saved_model_local_path, num_labels=2)

# 对每个用户的帖子集合进行分类
user_agg["user_label"] = user_agg.user_posts_cat.apply(
    lambda posts: label_cls(posts, tokenizer, model)
)
# 转换为机器人/人类标签
user_agg["user_label"] = np.where(user_agg.user_label == "LABEL_1", "bot", "human")

方案二:基于Accelerate的自动多GPU分配(模型并行,适合大模型)

如果你的模型较大,单GPU无法容纳,可以使用transformers的accelerate库实现自动模型层分配:

1. 安装依赖

pip install accelerate

2. 修改模型加载与推理代码

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer, pipeline
from collections import Counter
import pandas as pd
import numpy as np

def load_nlp(directory_path, num_labels=2):
    tokenizer_ = AutoTokenizer.from_pretrained(directory_path)
    # 使用device_map="auto"自动将模型层分配到多GPU
    model_ = AutoModelForSequenceClassification.from_pretrained(
        directory_path, num_labels=num_labels, device_map="auto"
    )
    pipe = pipeline("text-classification", model=model_, tokenizer=tokenizer_)
    return pipe

saved_model_local_path = "model_path"
pipe = load_nlp(saved_model_local_path, num_labels=2)

def label_cls(posts, pipe=pipe) -> str:
    labels = pipe(posts, padding=True, truncation=True, max_length=512)
    most_common_label = Counter([item["label"] for item in labels]).most_common(1)[0][0]
    return most_common_label

user_agg["user_label"] = [label_cls(posts) for posts in user_agg.user_posts_cat]
user_agg["user_label"] = np.where(user_agg.user_label == "LABEL_1", "bot", "human")

说明

  • 方案一为数据并行:将同一份模型复制到所有GPU,每个GPU处理一部分批量数据,适合小模型加速批量推理任务。
  • 方案二为模型并行:将模型的不同层分配到不同GPU,适合单个GPU无法容纳的大模型。

内容的提问来源于stack exchange,提问作者Kevin Li

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 11:37:07