在Google Colab加载LLaMa 2 70B时AssertionError的解决方法
问题:Google Colab加载Llama-2-70B-Chat-GGML模型触发AssertionError
问题背景
在Google Colab中尝试加载TheBloke/Llama-2-70B-Chat-GGML的q4_0量化模型,此前成功运行过13B版本,但加载70B模型时触发AssertionError,报错指向assert self.model is not None,说明模型加载失败。
运行代码
!pip install huggingface_hub model_name_or_path = "TheBloke/Llama-2-70B-Chat-GGML" model_basename = "llama-2-70b-chat.ggmlv3.q4_0.bin" from huggingface_hub import hf_hub_download from llama_cpp import Llama model_path = hf_hub_download(repo_id=model_name_or_path, filename=model_basename) # GPU lcpp_llm = None lcpp_llm = Llama( model_path=model_path, n_threads=2, # CPU cores n_batch=512, # Should be between 1 and n_ctx, consider the amount of VRAM in your GPU. n_gpu_layers=32 # Change this value based on your model and your GPU VRAM pool. )
报错信息
AssertionError Traceback (most recent call last) <ipython-input-51-da96b2fa6a04> in <cell line: 3>() 1 # GPU 2 lcpp_llm = None ----> 3 lcpp_llm = Llama( 4 model_path=model_path, 5 n_threads=2, # CPU cores /usr/local/lib/python3.10/dist-packages/llama_cpp/llama.py in __init__(self, model_path, n_ctx, n_parts, n_gpu_layers, seed, f16_kv, logits_all, vocab_only, use_mmap, use_mlock, embedding, n_threads, n_batch, last_n_tokens_size, lora_base, lora_path, low_vram, tensor_split, rope_freq_base, rope_freq_scale, n_gqa, rms_norm_eps, verbose) 311 self.model_path.encode("utf-8"), self.params 312 ) --> 313 assert self.model is not None 314 315 self.ctx = llama_cpp.llama_new_context_with_model(self.model, self.params) AssertionError:
原因分析
- 显存不足:Llama-2-70B q4_0量化后约占35GB显存,Colab免费版T4仅16GB,无法容纳足够的GPU层,导致模型加载失败。
- llama-cpp-python未启用CUDA优化:默认pip安装的版本可能未编译CUDA支持,无法利用GPU加速,甚至导致加载失败。
- 参数设置不合理:
n_gpu_layers设置过低或过高,n_threads、n_batch未适配Colab环境,影响加载效率。
修复方案
1. 确认GPU配置
必须使用Colab Pro/Pro+的A100(40GB/80GB显存),免费版无法运行70B模型。运行以下代码查看当前GPU:
!nvidia-smi
2. 重新安装带CUDA优化的llama-cpp-python
执行以下命令强制安装支持CUDA的版本,避免兼容性问题:
!CMAKE_ARGS="-DLLAMA_CUBLAS=on" FORCE_CMAKE=1 pip install llama-cpp-python==0.2.24 --force-reinstall --upgrade
3. 调整模型加载参数
适配A100 40GB显存的参数设置,最大化利用GPU资源:
!pip install huggingface_hub model_name_or_path = "TheBloke/Llama-2-70B-Chat-GGML" model_basename = "llama-2-70b-chat.ggmlv3.q4_0.bin" from huggingface_hub import hf_hub_download from llama_cpp import Llama model_path = hf_hub_download(repo_id=model_name_or_path, filename=model_basename) # 适配A100 40GB显存的参数配置 lcpp_llm = Llama( model_path=model_path, n_threads=8, # 利用Colab的8核CPU线程 n_batch=1024, # 增大批次提升加载与生成效率 n_gpu_layers=80, # 将70B模型的所有80层移至GPU n_ctx=2048, # 上下文窗口,可根据需求调整 verbose=True # 开启日志,便于排查加载问题 )
4. 验证模型运行
添加测试代码确认模型正常工作:
prompt = "用简单的语言解释量子计算:" output = lcpp_llm( prompt=prompt, max_tokens=200, temperature=0.7, top_p=0.9, stop=["</s>"], echo=False ) print(output["choices"][0]["text"])
备选方案(无Pro权限)
若无法使用A100,可尝试q2_k量化版本(约15GB),但生成效果会下降,且需调整参数适配免费T4:
model_basename = "llama-2-70b-chat.ggmlv3.q2_k.bin" # 加载参数调整 lcpp_llm = Llama( model_path=model_path, n_threads=8, n_batch=512, n_gpu_layers=20, # 仅将部分层移至GPU,剩余用CPU运行 verbose=True )
注:CPU运行70B模型速度极慢,仅作应急测试用。
内容的提问来源于stack exchange,提问作者Hoang Cuong Nguyen
相关产品推荐
相关产品推荐

