You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mac M1 Max上用LoRA微调模型时Transformer训练函数报Device()错误

问题分析与解决方案

核心问题

你遇到的错误源于BitsAndBytes的4bit量化(QLoRA)目前不支持Apple Silicon的MPS设备,同时代码中device_map="mps"与QLoRA配置的冲突,导致Hugging Face Accelerator在准备模型和优化器时出现设备参数异常。


具体解决方法

方法1:放弃4bit量化,用半精度在MPS上微调

如果你的模型大小适配M1 Max显存(比如7B模型半精度约占14GB显存,24GB版本M1 Max可支持),直接移除BitsAndBytes配置,改用原生MPS加速:

# 移除所有bnb_config相关代码
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.float16,  # 半精度减少显存占用
    device_map="mps",
)

# 训练参数开启半精度加速
training_arguments = TrainingArguments(
    # 其他参数保持不变
    fp16=True,
    device="mps"
)

此时LoRA配置保持正常即可,无需QLoRA相关设置。

方法2:切换到CPU运行(兼容QLoRA但速度慢)

强制模型在CPU上加载和训练,规避MPS与BitsAndBytes的兼容性问题:

device_map = "cpu"
# 移除torch.set_default_device(device_map)

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    quantization_config=bnb_config,
    device_map=device_map,
)

training_arguments = TrainingArguments(
    # 其他参数保持不变
    device="cpu"
)

适合小模型测试,大模型训练速度会显著降低。

方法3:尝试MPS兼容的8bit量化

部分模型支持Hugging Face原生的8bit量化(对MPS有限支持),可替代QLoRA:

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    load_in_8bit=True,
    device_map="mps",
)

注意:不是所有模型都支持MPS的8bit量化,需自行测试验证。


额外关键调整

  1. 移除torch.set_default_device(device_map),改为手动将模型和数据绑定到设备,避免加速器自动处理时的冲突:
# 加载模型后手动移到MPS
model = model.to("mps")

# 预处理时将数据张量移到MPS
def preprocess_function(examples):
    inputs = tokenizer(examples["text"], truncation=True, max_length=max_seq_length)
    inputs["labels"] = inputs["input_ids"].copy()
    for k, v in inputs.items():
        inputs[k] = torch.tensor(v).to("mps")
    return inputs

dataset = dataset.map(preprocess_function)
  1. 确保training_arguments中的设备设置与代码逻辑一致,避免冲突。

内容的提问来源于stack exchange,提问作者trialcritic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 11:31:22