使用Llama-2-13b-chat-hf时遇CUDA内核镜像不可用错误求解决
问题:运行Llama-2-13b-chat-hf模型时出现CUDA kernel错误
运行meta-llama/Llama-2-13b-chat-hf模型时触发以下错误:
RuntimeError: CUDA error: no kernel image is available for execution on the device CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1. Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
环境信息
NVIDIA驱动与CUDA版本
- nvidia-smi输出:
| NVIDIA-SMI 465.19.01 Driver Version: 465.19.01 CUDA Version: 11.3 |
- NVCC版本:
nvcc: NVIDIA (R) Cuda compiler driver Copyright (c) 2005-2022 NVIDIA Corporation Built on Wed_Sep_21_10:33:58_PDT_2022 Cuda compilation tools, release 11.8, V11.8.89 Build cuda_11.8.r11.8/compiler.31833905_0
依赖包版本
- PyTorch相关:
torch==2.0.0+cu118 torchaudio==2.0.1+cu118 torchvision==0.15.1+cu118
- Transformers:
transformers==4.37.2
硬件与系统
- 硬件:多张NVIDIA GeForce RTX 2080 Ti + 1张NVIDIA GeForce GT 710
- 系统:Ubuntu 16
额外警告与验证信息
运行时收到警告:
Found GPU9 NVIDIA GeForce GT 710 which is of cuda capability 3.5. PyTorch no longer supports this GPU because it is too old. The minimum cuda capability supported by this library is 3.7.
bitsandbytes安装验证成功:
++++++++++++++++++++++++++ OTHER +++++++++++++++++++++++++++ COMPILED_WITH_CUDA = True COMPUTE_CAPABILITIES_PER_GPU = ['7.5', '7.5', '7.5', '7.5', '7.5', '7.5', '7.5', '7.5', '7.5', '3.5'] ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ ++++++++++++++++++++++ DEBUG INFO END ++++++++++++++++++++++ ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ Running a quick check that: + library is importable + CUDA function is callable WARNING: Please be sure to sanitize sensible info from any such env vars! SUCCESS! Installation was successful!
PyTorch测试代码:
import torch import sys print('A', sys.version) print('B', torch.__version__) print('C', torch.cuda.is_available()) print('D', torch.backends.cudnn.enabled) device = torch.device('cuda') print('E', torch.cuda.get_device_properties(device)) print('F', torch.tensor([1.0, 2.0]).cuda())
测试输出:
A 3.11.7 (main, Dec 15 2023, 18:12:31) [GCC 11.2.0] B 2.0.0+cu118 C True D True UserWarning: Found GPU9 NVIDIA GeForce GT 710 which is of cuda capability 3.5. PyTorch no longer supports this GPU because it is too old. The minimum cuda capability supported by this library is 3.7. warnings.warn(old_gpu_warn % (d, name, major, minor, min_arch // 10, min_arch % 10)) E _CudaDeviceProperties(name='NVIDIA GeForce RTX 2080 Ti', major=7, minor=5, total_memory=11019MB, multi_processor_count=68) F tensor([1., 2.], device='cuda:0')
解决方案
核心原因
GT 710的CUDA计算能力为3.5,而PyTorch 2.0.0+cu118最低要求3.7。即使测试代码默认使用RTX 2080 Ti,但PyTorch初始化时会检测所有GPU,部分库(如bitsandbytes、模型CUDA kernel)可能尝试在不支持的GT 710上加载代码,导致no kernel image is available错误。
具体解决步骤
- 屏蔽GT 710显卡:运行代码前设置环境变量,让PyTorch仅识别可用的RTX 2080 Ti。先通过
nvidia-smi确认GT 710的GPU编号,将其排除在外:
或在Python代码开头添加:export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7,8import os os.environ["CUDA_VISIBLE_DEVICES"] = "0,1,2,3,4,5,6,7,8" - 升级NVIDIA驱动:当前驱动465.19.01仅支持CUDA 11.3,而PyTorch使用的是CUDA 11.8,建议升级驱动到≥450.80.02(推荐470.x及以上版本),解决驱动与CUDA版本不匹配的潜在问题。
- 明确指定模型加载设备:加载模型时指定使用支持的GPU,避免自动分配到GT 710:
from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained( "meta-llama/Llama-2-13b-chat-hf", device_map="auto", torch_dtype=torch.float16 ) # 或手动指定设备 device = torch.device("cuda:0") model = model.to(device)
内容的提问来源于stack exchange,提问作者bbqman
相关产品推荐
相关产品推荐

