You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Google Colab加载LLaMa 2 70B时AssertionError的解决方法

问题:Google Colab加载Llama-2-70B-Chat-GGML模型触发AssertionError

问题背景

在Google Colab中尝试加载TheBloke/Llama-2-70B-Chat-GGML的q4_0量化模型,此前成功运行过13B版本,但加载70B模型时触发AssertionError,报错指向assert self.model is not None,说明模型加载失败。

运行代码

!pip install huggingface_hub
model_name_or_path = "TheBloke/Llama-2-70B-Chat-GGML"
model_basename = "llama-2-70b-chat.ggmlv3.q4_0.bin"

from huggingface_hub import hf_hub_download
from llama_cpp import Llama

model_path = hf_hub_download(repo_id=model_name_or_path, filename=model_basename)

# GPU
lcpp_llm = None
lcpp_llm = Llama(
    model_path=model_path,
    n_threads=2, # CPU cores
    n_batch=512, # Should be between 1 and n_ctx, consider the amount of VRAM in your GPU.
    n_gpu_layers=32 # Change this value based on your model and your GPU VRAM pool.
    )

报错信息

AssertionError                            Traceback (most recent call last)
<ipython-input-51-da96b2fa6a04> in <cell line: 3>()
      1 # GPU
      2 lcpp_llm = None
----> 3 lcpp_llm = Llama(
      4     model_path=model_path,
      5     n_threads=2, # CPU cores

/usr/local/lib/python3.10/dist-packages/llama_cpp/llama.py in __init__(self, model_path, n_ctx, n_parts, n_gpu_layers, seed, f16_kv, logits_all, vocab_only, use_mmap, use_mlock, embedding, n_threads, n_batch, last_n_tokens_size, lora_base, lora_path, low_vram, tensor_split, rope_freq_base, rope_freq_scale, n_gqa, rms_norm_eps, verbose)
    311             self.model_path.encode("utf-8"), self.params
    312         )
--> 313         assert self.model is not None
    314 
    315         self.ctx = llama_cpp.llama_new_context_with_model(self.model, self.params)

AssertionError: 

原因分析

  1. 显存不足:Llama-2-70B q4_0量化后约占35GB显存,Colab免费版T4仅16GB,无法容纳足够的GPU层,导致模型加载失败。
  2. llama-cpp-python未启用CUDA优化:默认pip安装的版本可能未编译CUDA支持,无法利用GPU加速,甚至导致加载失败。
  3. 参数设置不合理:n_gpu_layers设置过低或过高,n_threads、n_batch未适配Colab环境,影响加载效率。

修复方案

1. 确认GPU配置

必须使用Colab Pro/Pro+的A100(40GB/80GB显存),免费版无法运行70B模型。运行以下代码查看当前GPU:

!nvidia-smi

2. 重新安装带CUDA优化的llama-cpp-python

执行以下命令强制安装支持CUDA的版本,避免兼容性问题:

!CMAKE_ARGS="-DLLAMA_CUBLAS=on" FORCE_CMAKE=1 pip install llama-cpp-python==0.2.24 --force-reinstall --upgrade

3. 调整模型加载参数

适配A100 40GB显存的参数设置,最大化利用GPU资源:

!pip install huggingface_hub
model_name_or_path = "TheBloke/Llama-2-70B-Chat-GGML"
model_basename = "llama-2-70b-chat.ggmlv3.q4_0.bin"

from huggingface_hub import hf_hub_download
from llama_cpp import Llama

model_path = hf_hub_download(repo_id=model_name_or_path, filename=model_basename)

# 适配A100 40GB显存的参数配置
lcpp_llm = Llama(
    model_path=model_path,
    n_threads=8,  # 利用Colab的8核CPU线程
    n_batch=1024, # 增大批次提升加载与生成效率
    n_gpu_layers=80, # 将70B模型的所有80层移至GPU
    n_ctx=2048,   # 上下文窗口,可根据需求调整
    verbose=True  # 开启日志,便于排查加载问题
)

4. 验证模型运行

添加测试代码确认模型正常工作:

prompt = "用简单的语言解释量子计算:"
output = lcpp_llm(
    prompt=prompt,
    max_tokens=200,
    temperature=0.7,
    top_p=0.9,
    stop=["</s>"],
    echo=False
)
print(output["choices"][0]["text"])

备选方案(无Pro权限)

若无法使用A100,可尝试q2_k量化版本(约15GB),但生成效果会下降,且需调整参数适配免费T4:

model_basename = "llama-2-70b-chat.ggmlv3.q2_k.bin"
# 加载参数调整
lcpp_llm = Llama(
    model_path=model_path,
    n_threads=8,
    n_batch=512,
    n_gpu_layers=20, # 仅将部分层移至GPU,剩余用CPU运行
    verbose=True
)

注:CPU运行70B模型速度极慢,仅作应急测试用。

内容的提问来源于stack exchange,提问作者Hoang Cuong Nguyen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 23:10:37