You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LLAMA-13B推理触发AssertionError报错求助及解决方案

AssertionError 初始化Vigogne-2-13B-Instruct-GGML模型时的解决方案

在使用llama-cpp-python加载Vigogne-2-13B-Instruct-GGML模型推理时,触发如下断言错误:

---------------------------------------------------------------------------
AssertionError                            Traceback (most recent call last)
<ipython-input-5-7820a34f7358> in <cell line: 3>()
      1 # GPU
      2 lcpp_llm = None
----> 3 lcpp_llm = Llama(
      4     model_path=model_path,
      5     # n_gqa = 8,

/usr/local/lib/python3.10/dist-packages/llama_cpp/llama.py in __init__(self, model_path, n_ctx, n_parts, n_gpu_layers, seed, f16_kv, logits_all, vocab_only, use_mmap, use_mlock, embedding, n_threads, n_batch, last_n_tokens_size, lora_base, lora_path, low_vram, tensor_split, rope_freq_base, rope_freq_scale, n_gqa, rms_norm_eps, mul_mat_q, verbose)
    321                     self.model_path.encode("utf-8"), self.params
    322                 )
--> 323         assert self.model is not None
    324 
    325         if verbose:

AssertionError: 

相关代码片段

初始化Llama模型的代码:

# GPU
lcpp_llm = None
lcpp_llm = Llama(
    model_path=model_path,
    # n_gqa = 8,
    n_threads=2, # CPU cores,
    n_ctx = 4096,
    n_batch=512, # Should be between 1 and n_ctx, consider the amount of VRAM in your GPU.
    n_gpu_layers=32 # Change this value based on your model and your GPU VRAM pool.
)

model_path的定义:

model_path = hf_hub_download(repo_id=model_name_or_path, filename=model_basename)

使用的模型信息:

model_name_or_path = "TheBloke/Vigogne-2-13B-Instruct-GGML"
model_basename = "vigogne-2-13b-instruct.ggmlv3.q5_1.bin" # the model is in bin format

解决方案

将llama-cpp-python版本降级到v0.1.78即可解决该问题。

内容的提问来源于stack exchange,提问作者Malik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 14:22:20