You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决LLM微调报错:如何支持或规避device_map='auto'?

问题:Bio_ClinicalBERT微调时触发device_map不支持错误

在Google Colab T4 GPU环境下,导入emilyalsentzer/Bio_ClinicalBERT并基于医疗数据微调时,持续出现以下错误:

ValueError: BertLMHeadModel does not support device_map='auto'. To implement support, the model class needs to implement the _no_split_modules attribute.

相关代码

from transformers import AutoModelForSequenceClassification, AutoTokenizer

# Choose a model appropriate for your task
model_name = "emilyalsentzer/Bio_ClinicalBERT"
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Set device manually
device = "cuda" if torch.cuda.is_available() else "cpu"

# Load the model and move it to the selected device
model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels=2)
model.to(device)

# Move inputs to the same device
inputs = tokenizer("Your clinical note here", return_tensors="pt")
inputs = {k: v.to(device) for k, v in inputs.items()}

# Inference
model.eval()
with torch.no_grad():
    outputs = model(**inputs)

df['num_characters_input'] = df['input'].apply(lambda x: len(x))

model = None
clear_cache()
df_train = df
qlora_fine_tuning_config = yaml.safe_load(
"""
model_type: llm
base_model: emilyalsentzer/Bio_ClinicalBERT

input_features:
  - name: instruction
    type: text

output_features:
  - name: output
    type: text

prompt:
  template: >-
    Below is an instruction that describes a task, paired with an input
    that may provide further context. Write a response that appropriately
    completes the request.

    ### Instruction: {instruction}

    ### Input: {input}

    ### Response:

generation:
  temperature: .1
  max_new_tokens: 10

adapter:
  type: lora
  r: 8

quantization:
  bits: 8

trainer:
  type: finetune
  epochs: 4
  batch_size: 16
  eval_batch_size: 16
  gradient_accumulation_steps: 16
  learning_rate: 0.00001
  optimizer:
    type: adam
    params:
      eps: 1.e-8
      betas:
        - 0.9
        - 0.999
      weight_decay: 0
  learning_rate_scheduler:
    warmup_fraction: 0.06
    reduce_on_plateau: 1
"""
)


model = LudwigModel(config=qlora_fine_tuning_config, logging_level=logging.INFO)

results = model.train(dataset=df_train)

解决建议

1. 修改Ludwig配置,禁用自动设备映射

错误根源是Ludwig加载LLM时默认使用device_map='auto',但BERT类模型不支持该参数。在配置的base_model字段下添加model_config,强制关闭自动设备映射:

base_model: emilyalsentzer/Bio_ClinicalBERT
model_config:
  device_map: None

修改后的配置会让Ludwig使用你手动指定的GPU设备,避免触发自动映射逻辑。

2. 改用Hugging Face原生Trainer API(更灵活)

如果不需要依赖Ludwig的封装,可以直接用Hugging Face + PEFT实现QLoRA微调,完全控制设备分配:

import torch
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, Trainer, DataCollatorForLanguageModeling

model_name = "emilyalsentzer/Bio_ClinicalBERT"
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token

# 手动加载模型并指定设备
device = "cuda" if torch.cuda.is_available() else "cpu"
model = AutoModelForCausalLM.from_pretrained(model_name, device_map=None)
model = model.to(device)

# 配置QLoRA
lora_config = LoraConfig(
    r=8,
    lora_alpha=32,
    target_modules=["query", "key"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM"
)
model = get_peft_model(model, lora_config)

# 训练配置
training_args = TrainingArguments(
    output_dir="./bio_clinicalbert_qlora",
    per_device_train_batch_size=16,
    gradient_accumulation_steps=16,
    learning_rate=1e-5,
    num_train_epochs=4,
    logging_steps=10,
    fp16=True,
    optim="adamw_torch"
)

# 数据处理(需根据你的数据集调整)
def preprocess_function(examples):
    return tokenizer(examples["input"], truncation=True, max_length=512)

tokenized_dataset = df_train.map(preprocess_function, batched=True)
data_collator = DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False)

# 启动训练
trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_dataset,
    data_collator=data_collator
)
trainer.train()

3. 升级依赖库

旧版本的transformers、peft或ludwig可能存在设备映射的兼容性问题,执行以下命令升级:

pip install --upgrade transformers peft ludwig torch

内容的提问来源于stack exchange,提问作者pctree

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 12:57:07