基于Vicuna+LangChain+LlamaIndex搭建自托管LLM遇CUDA内存不足问题
问题
我希望搭建一个可接入自定义数据(Slack对话)的自托管LLM模型,选择Vicuna作为ChatGPT替代方案,编写了如下代码:
from llama_index import SimpleDirectoryReader, LangchainEmbedding, GPTListIndex, \ GPTSimpleVectorIndex, PromptHelper, LLMPredictor, Document, ServiceContext from langchain.embeddings.huggingface import HuggingFaceEmbeddings import torch from langchain.llms.base import LLM from transformers import pipeline, AutoTokenizer, AutoModelForCausalLM !export PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512 class CustomLLM(LLM): model_name = "eachadea/vicuna-13b-1.1" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained(model_name) pipeline = pipeline("text2text-generation", model=model, tokenizer=tokenizer, device=0, model_kwargs={"torch_dtype":torch.bfloat16}) def _call(self, prompt, stop=None): return self.pipeline(prompt, max_length=9999)[0]["generated_text"] def _identifying_params(self): return {"name_of_model": self.model_name} def _llm_type(self): return "custom" llm_predictor = LLMPredictor(llm=CustomLLM())
运行时出现CUDA内存不足错误:
OutOfMemoryError: CUDA out of memory. Tried to allocate 270.00 MiB (GPU 0; 22.03 GiB total capacity; 21.65 GiB already allocated; 94.88 MiB free; 21.65 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF
运行前!nvidia-smi输出:
Thu Apr 20 18:04:00 2023 +---------------------------------------------------------------------------------------+ | NVIDIA-SMI 530.30.02 Driver Version: 530.30.02 CUDA Version: 12.1 | |-----------------------------------------+----------------------+----------------------+ | GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+======================+======================| | 0 NVIDIA A10G Off| 00000000:00:1E.0 Off | 0 | | 0% 23C P0 52W / 300W| 0MiB / 23028MiB | 18% Default | | | | N/A | +-----------------------------------------+----------------------+----------------------+ +---------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=======================================================================================| | No running processes found | +---------------------------------------------------------------------------------------+
解决方案
针对Vicuna-13B在A10G(22GB显存)上的OOM问题,可通过以下代码调整解决:
启用4-bit量化:借助
bitsandbytes库实现模型量化,大幅降低显存占用。
先安装依赖:pip install bitsandbytes accelerate修改后的完整代码:
import os import torch from llama_index import LLMPredictor from langchain.llms.base import LLM from transformers import pipeline, AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig # 正确设置显存配置环境变量 os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:512" torch.cuda.empty_cache() # 配置4-bit量化参数 bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16 ) class CustomLLM(LLM): model_name = "eachadea/vicuna-13b-1.1" tokenizer = AutoTokenizer.from_pretrained(model_name) # 加载量化后的模型,自动分配设备 model = AutoModelForCausalLM.from_pretrained( model_name, quantization_config=bnb_config, device_map="auto", torch_dtype=torch.bfloat16 ) # Vicuna是自回归模型,任务类型改为text-generation pipeline = pipeline("text-generation", model=model, tokenizer=tokenizer, device_map="auto", model_kwargs={"torch_dtype":torch.bfloat16}) def _call(self, prompt, stop=None): # 限制生成新token数量,避免显存溢出 outputs = self.pipeline(prompt, max_new_tokens=512, do_sample=True, temperature=0.7) return outputs[0]["generated_text"].replace(prompt, "") def _identifying_params(self): return {"name_of_model": self.model_name} def _llm_type(self): return "custom" llm_predictor = LLMPredictor(llm=CustomLLM())关键调整说明:
- 替换
text2text-generation为text-generation,匹配Vicuna模型的任务类型 - 使用
max_new_tokens替代max_length,避免计算冗余上下文长度 - 加载模型时添加
device_map="auto",让框架自动分配模型层到GPU/CPU - 代码开头添加
torch.cuda.empty_cache()清理显存碎片
- 替换
内容的提问来源于stack exchange,提问作者Ben
相关产品推荐
相关产品推荐

