You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在cuda:1上微调Llama2时设备不匹配问题求助

解决Llama2 4位量化微调时设备不匹配的ValueError问题

问题核心

你遇到的错误本质是模型加载在cuda:1,但训练进程默认使用了其他设备(通常是cuda:0),导致4位量化模型与训练张量的设备不一致。即便指定了device_map={'':1},训练器、LoRA层或数据张量仍可能被分配到默认设备,引发冲突。

具体解决步骤

以下是针对性的代码修改方案:

  1. 全局锁定CUDA设备
    在代码开头显式设置当前使用的CUDA设备为1,确保后续所有操作(模型加载、数据处理、训练)都默认指向该设备:
import torch
torch.cuda.set_device(1)  # 强制所有CUDA操作默认使用cuda:1
  1. 明确指定device_map为torch.device对象
    将device_map的写法从{"": 1}改为更明确的{"": torch.device("cuda:1")},避免设备索引解析歧义:
base_model = AutoModelForCausalLM.from_pretrained(
    base_model_name,
    quantization_config=bnb_config,
    device_map={"": torch.device("cuda:1")},  # 修改此处
    trust_remote_code=True,
    use_auth_token=True,
)
  1. 修正training_args的设备配置
    确保你的training_args中明确指定训练设备为cuda:1,并开启与模型匹配的精度设置:
from transformers import TrainingArguments

training_args = TrainingArguments(
    # 其他原有参数保持不变
    device="cuda:1",  # 指定训练设备
    bf16=True,  # 和模型的bnb_4bit_compute_dtype=torch.bfloat16匹配
    use_cpu=False,  # 禁用CPU训练
)
  1. 验证模型设备分配
    加载模型后添加验证代码,确认模型确实在cuda:1上:
print("模型所在设备:", next(base_model.parameters()).device)
# 正常输出应为: 模型所在设备: cuda:1

修改后的完整代码示例

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig, TrainingArguments
from peft import LoraConfig
from trl import SFTTrainer

# 全局指定CUDA设备
torch.cuda.set_device(1)

# 加载4位量化基础模型
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
)
    
base_model = AutoModelForCausalLM.from_pretrained(
    base_model_name,
    quantization_config=bnb_config,
    device_map={"": torch.device("cuda:1")},
    trust_remote_code=True,
    use_auth_token=True,
)
# 验证设备分配
print("模型所在设备:", next(base_model.parameters()).device)

base_model.config.use_cache = False
tokenizer = AutoTokenizer.from_pretrained(base_model_name, use_auth_token=True)
    
# 配置LoRA层
peft_config = LoraConfig(
    r=16,
    lora_alpha=64,
    lora_dropout=0.1,
    target_modules=["q_proj", "v_proj"],
    bias="none",
    task_type="CAUSAL_LM",
)

# 修正训练参数的设备配置
training_args = TrainingArguments(
    output_dir="./results",
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    learning_rate=2e-4,
    num_train_epochs=3,
    logging_steps=10,
    save_strategy="epoch",
    device="cuda:1",
    bf16=True,
    use_cpu=False,
)
    
trainer = SFTTrainer(
    model=base_model,
    train_dataset=dataset,
    peft_config=peft_config,
    packing=True,
    max_seq_length=None,
    dataset_text_field="text",
    tokenizer=tokenizer,
    args=training_args,
)
trainer.train()

关键说明

4位/8位量化模型对设备一致性要求极高,模型、LoRA层、训练数据、优化器张量必须全部在同一设备上,否则会触发设备不匹配错误。全局设置torch.cuda.set_device(1)是最稳妥的方式,能避免后续代码中隐性的设备分配问题。

内容的提问来源于stack exchange,提问作者user1564762

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 07:37:10