You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用QLoRA与Peft微调Gemma模型时遭遇CUDA设备错误求助

解决Gemma-7b QLoRA微调时CUDA设备断言失败问题

问题场景

尝试用Peft+QLoRA微调google/gemma-7b模型,昨日测试微调流程正常,今日加载模型时触发CUDA相关断言错误,核心提示设备索引无效。

复现代码

model_id = "google/gemma-7b"

bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16
)

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, 
quantization_config=bnb_config, 
device_map={0:""})

#model.gradient_checkpointing_enable()

train_dataset, val_dataset, data_collator = load_dataset(train_data_path, val_data_path, tokenizer)

核心报错信息

RuntimeError: device >= 0 && device < num_gpus INTERNAL ASSERT FAILED at "../aten/src/ATen/cuda/CUDAContext.cpp":50, please report a bug to PyTorch. device=1, num_gpus=

DeferredCudaCallError: CUDA call failed lazily at initialization with error: device >= 0 && device < num_gpus INTERNAL ASSERT FAILED at "../aten/src/ATen/cuda/CUDAContext.cpp":50, please report a bug to PyTorch. device=1, num_gpus=

RuntimeError: Failed to import transformers.integrations.bitsandbytes because of the following error (look up to see its traceback):
CUDA call failed lazily at initialization with error: device >= 0 && device < num_gpus INTERNAL ASSERT FAILED at "../aten/src/ATen/cuda/CUDAContext.cpp":50, please report a bug to PyTorch. device=1, num_gpus=

解决方法

1. 确认CUDA设备识别状态

运行以下代码检查GPU是否被正确识别:

import torch
print(torch.cuda.is_available())
print(torch.cuda.device_count())
  • 若返回False或0,说明GPU未加载,需重启运行环境、检查显卡驱动/CUDA toolkit是否正常;云平台实例则优先考虑重启实例或联系运维确认GPU挂载状态。

2. 修正device_map配置

代码中device_map={0:""}写法不规范,替换为以下两种方式之一:

  • 自动分配设备:device_map="auto"
  • 明确指定第0块GPU:device_map={"":0}

3. 重置CUDA上下文

  • 先清空缓存:torch.cuda.empty_cache()
  • 重启Python内核后重新加载模型
  • 若仍报错,在代码开头添加torch.cuda.set_device(0)强制指定使用第0块GPU

4. 检查依赖版本兼容性

确认bitsandbytes与PyTorch、CUDA版本匹配,重新安装对应版本的bitsandbytes:

pip install bitsandbytes==0.42.0  # 适配PyTorch 2.0+版本,可根据实际调整

5. 验证CUDA环境变量

检查CUDA_VISIBLE_DEVICES是否指向有效GPU:

  • Linux/macOS:运行echo $CUDA_VISIBLE_DEVICES
  • Windows:运行echo %CUDA_VISIBLE_DEVICES%
  • 若输出为空或无效索引,设置为export CUDA_VISIBLE_DEVICES=0(Linux)或set CUDA_VISIBLE_DEVICES=0(Windows)

内容的提问来源于stack exchange,提问作者eneko valero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 06:03:14