在cuda:1上微调Llama2时设备不匹配问题求助
解决Llama2 4位量化微调时设备不匹配的ValueError问题
问题核心
你遇到的错误本质是模型加载在cuda:1,但训练进程默认使用了其他设备(通常是cuda:0),导致4位量化模型与训练张量的设备不一致。即便指定了device_map={'':1},训练器、LoRA层或数据张量仍可能被分配到默认设备,引发冲突。
具体解决步骤
以下是针对性的代码修改方案:
- 全局锁定CUDA设备
在代码开头显式设置当前使用的CUDA设备为1,确保后续所有操作(模型加载、数据处理、训练)都默认指向该设备:
import torch torch.cuda.set_device(1) # 强制所有CUDA操作默认使用cuda:1
- 明确指定device_map为torch.device对象
将device_map的写法从{"": 1}改为更明确的{"": torch.device("cuda:1")},避免设备索引解析歧义:
base_model = AutoModelForCausalLM.from_pretrained( base_model_name, quantization_config=bnb_config, device_map={"": torch.device("cuda:1")}, # 修改此处 trust_remote_code=True, use_auth_token=True, )
- 修正training_args的设备配置
确保你的training_args中明确指定训练设备为cuda:1,并开启与模型匹配的精度设置:
from transformers import TrainingArguments training_args = TrainingArguments( # 其他原有参数保持不变 device="cuda:1", # 指定训练设备 bf16=True, # 和模型的bnb_4bit_compute_dtype=torch.bfloat16匹配 use_cpu=False, # 禁用CPU训练 )
- 验证模型设备分配
加载模型后添加验证代码,确认模型确实在cuda:1上:
print("模型所在设备:", next(base_model.parameters()).device) # 正常输出应为: 模型所在设备: cuda:1
修改后的完整代码示例
import torch from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig, TrainingArguments from peft import LoraConfig from trl import SFTTrainer # 全局指定CUDA设备 torch.cuda.set_device(1) # 加载4位量化基础模型 bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, ) base_model = AutoModelForCausalLM.from_pretrained( base_model_name, quantization_config=bnb_config, device_map={"": torch.device("cuda:1")}, trust_remote_code=True, use_auth_token=True, ) # 验证设备分配 print("模型所在设备:", next(base_model.parameters()).device) base_model.config.use_cache = False tokenizer = AutoTokenizer.from_pretrained(base_model_name, use_auth_token=True) # 配置LoRA层 peft_config = LoraConfig( r=16, lora_alpha=64, lora_dropout=0.1, target_modules=["q_proj", "v_proj"], bias="none", task_type="CAUSAL_LM", ) # 修正训练参数的设备配置 training_args = TrainingArguments( output_dir="./results", per_device_train_batch_size=4, gradient_accumulation_steps=4, learning_rate=2e-4, num_train_epochs=3, logging_steps=10, save_strategy="epoch", device="cuda:1", bf16=True, use_cpu=False, ) trainer = SFTTrainer( model=base_model, train_dataset=dataset, peft_config=peft_config, packing=True, max_seq_length=None, dataset_text_field="text", tokenizer=tokenizer, args=training_args, ) trainer.train()
关键说明
4位/8位量化模型对设备一致性要求极高,模型、LoRA层、训练数据、优化器张量必须全部在同一设备上,否则会触发设备不匹配错误。全局设置torch.cuda.set_device(1)是最稳妥的方式,能避免后续代码中隐性的设备分配问题。
内容的提问来源于stack exchange,提问作者user1564762
相关产品推荐
相关产品推荐

