You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

启用llama_cublas后运行模型仅占用75MB VRAM的问题咨询

问题描述

我已启用llama_cublas适配NVIDIA CUDA Toolkit,执行命令make LLAMA_CUBLAS=1后编译成功。但运行模型时,通过nvidia-smi监控显存占用,发现仅使用了75MB VRAM,模型未加载至GPU。相关运行日志及nvidia-smi输出如下:

运行日志

llm_load_tensors: using CUDA for GPU acceleration
llm_load_tensors: mem required  = 13189.99 MB
llm_load_tensors: offloading 0 repeating layers to GPU
llm_load_tensors: offloaded 0/43 layers to GPU
llm_load_tensors: VRAM used: 0.00 MB
....................................................................................................
llama_new_context_with_model: n_ctx      = 512
llama_new_context_with_model: freq_base  = 10000.0
llama_new_context_with_model: freq_scale = 1
llama_new_context_with_model: kv self size  =  400.00 MB
llama_new_context_with_model: compute buffer total size = 81.13 MB
llama_new_context_with_model: VRAM scratch buffer: 75.00 MB
llama_new_context_with_model: total VRAM used: 75.00 MB (model: 0.00 MB, context: 75.00 MB)

nvidia-smi输出

Tue Oct 24 10:53:17 2023       
+---------------------------------------------------------------------------------------+
| NVIDIA-SMI 535.113.01             Driver Version: 535.113.01   CUDA Version: 12.2     |
|-----------------------------------------+----------------------+----------------------+
| GPU  Name                 Persistence-M | Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |         Memory-Usage | GPU-Util  Compute M. |
|                                         |                      |               MIG M. |
|=========================================+======================+======================|
|   0  NVIDIA GeForce RTX 4050 ...    Off | 00000000:01:00.0 Off |                  N/A |
| N/A   42C    P8               5W /  30W |      89MiB /  6141MiB |      0%      Default |
|                                         |                      |                  N/A |
+-----------------------------------------+----------------------+----------------------+
                                                                                          
+---------------------------------------------------------------------------------------+
| Processes:                                                                            |
|  GPU   GI   CI        PID   Type   Process name                            GPU Memory |
|        ID   ID                                                             Usage      |
|=======================================================================================|
|    0   N/A  N/A      1991      G   /usr/lib/xorg/Xorg                            4MiB |
+---------------------------------------------------------------------------------------+
解决方法

从日志里的offloaded 0/43 layers to GPU可知,没有任何模型层被卸载到GPU,核心问题是运行模型时未指定GPU卸载参数:

  • 运行模型时必须显式添加--n-gpu-layers参数,设置要卸载到GPU的层数。比如模型总共有43层,可尝试--n-gpu-layers 43;如果GPU显存不足(你的RTX4050只有约6GB,模型需要13GB),可以设置--n-gpu-layers -1让程序自动适配可卸载的最大层数,或者手动指定合理数值,比如--n-gpu-layers 20。
  • 确认运行命令调用的是编译好的支持CUDA的二进制文件,避免误启动未开启CUDA支持的版本。

内容的提问来源于stack exchange,提问作者djbritt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 20:38:32