解决LLM微调报错:如何支持或规避device_map='auto'?
问题:Bio_ClinicalBERT微调时触发device_map不支持错误
在Google Colab T4 GPU环境下,导入emilyalsentzer/Bio_ClinicalBERT并基于医疗数据微调时,持续出现以下错误:
ValueError: BertLMHeadModel does not support
device_map='auto'. To implement support, the model class needs to implement the_no_split_modulesattribute.
相关代码
from transformers import AutoModelForSequenceClassification, AutoTokenizer # Choose a model appropriate for your task model_name = "emilyalsentzer/Bio_ClinicalBERT" tokenizer = AutoTokenizer.from_pretrained(model_name) # Set device manually device = "cuda" if torch.cuda.is_available() else "cpu" # Load the model and move it to the selected device model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels=2) model.to(device) # Move inputs to the same device inputs = tokenizer("Your clinical note here", return_tensors="pt") inputs = {k: v.to(device) for k, v in inputs.items()} # Inference model.eval() with torch.no_grad(): outputs = model(**inputs) df['num_characters_input'] = df['input'].apply(lambda x: len(x)) model = None clear_cache() df_train = df qlora_fine_tuning_config = yaml.safe_load( """ model_type: llm base_model: emilyalsentzer/Bio_ClinicalBERT input_features: - name: instruction type: text output_features: - name: output type: text prompt: template: >- Below is an instruction that describes a task, paired with an input that may provide further context. Write a response that appropriately completes the request. ### Instruction: {instruction} ### Input: {input} ### Response: generation: temperature: .1 max_new_tokens: 10 adapter: type: lora r: 8 quantization: bits: 8 trainer: type: finetune epochs: 4 batch_size: 16 eval_batch_size: 16 gradient_accumulation_steps: 16 learning_rate: 0.00001 optimizer: type: adam params: eps: 1.e-8 betas: - 0.9 - 0.999 weight_decay: 0 learning_rate_scheduler: warmup_fraction: 0.06 reduce_on_plateau: 1 """ ) model = LudwigModel(config=qlora_fine_tuning_config, logging_level=logging.INFO) results = model.train(dataset=df_train)
解决建议
1. 修改Ludwig配置,禁用自动设备映射
错误根源是Ludwig加载LLM时默认使用device_map='auto',但BERT类模型不支持该参数。在配置的base_model字段下添加model_config,强制关闭自动设备映射:
base_model: emilyalsentzer/Bio_ClinicalBERT model_config: device_map: None
修改后的配置会让Ludwig使用你手动指定的GPU设备,避免触发自动映射逻辑。
2. 改用Hugging Face原生Trainer API(更灵活)
如果不需要依赖Ludwig的封装,可以直接用Hugging Face + PEFT实现QLoRA微调,完全控制设备分配:
import torch from peft import LoraConfig, get_peft_model from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments, Trainer, DataCollatorForLanguageModeling model_name = "emilyalsentzer/Bio_ClinicalBERT" tokenizer = AutoTokenizer.from_pretrained(model_name) tokenizer.pad_token = tokenizer.eos_token # 手动加载模型并指定设备 device = "cuda" if torch.cuda.is_available() else "cpu" model = AutoModelForCausalLM.from_pretrained(model_name, device_map=None) model = model.to(device) # 配置QLoRA lora_config = LoraConfig( r=8, lora_alpha=32, target_modules=["query", "key"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM" ) model = get_peft_model(model, lora_config) # 训练配置 training_args = TrainingArguments( output_dir="./bio_clinicalbert_qlora", per_device_train_batch_size=16, gradient_accumulation_steps=16, learning_rate=1e-5, num_train_epochs=4, logging_steps=10, fp16=True, optim="adamw_torch" ) # 数据处理(需根据你的数据集调整) def preprocess_function(examples): return tokenizer(examples["input"], truncation=True, max_length=512) tokenized_dataset = df_train.map(preprocess_function, batched=True) data_collator = DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False) # 启动训练 trainer = Trainer( model=model, args=training_args, train_dataset=tokenized_dataset, data_collator=data_collator ) trainer.train()
3. 升级依赖库
旧版本的transformers、peft或ludwig可能存在设备映射的兼容性问题,执行以下命令升级:
pip install --upgrade transformers peft ludwig torch
内容的提问来源于stack exchange,提问作者pctree
相关产品推荐
相关产品推荐

