You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Vicuna+LangChain+LlamaIndex搭建自托管LLM遇CUDA内存不足问题

问题

我希望搭建一个可接入自定义数据(Slack对话)的自托管LLM模型,选择Vicuna作为ChatGPT替代方案,编写了如下代码:

from llama_index import SimpleDirectoryReader, LangchainEmbedding, GPTListIndex, \
    GPTSimpleVectorIndex, PromptHelper, LLMPredictor, Document, ServiceContext
from langchain.embeddings.huggingface import HuggingFaceEmbeddings
import torch
from langchain.llms.base import LLM
from transformers import pipeline, AutoTokenizer, AutoModelForCausalLM

!export PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512
    
class CustomLLM(LLM):
    model_name = "eachadea/vicuna-13b-1.1"
    tokenizer = AutoTokenizer.from_pretrained(model_name)
    model = AutoModelForCausalLM.from_pretrained(model_name)

    pipeline = pipeline("text2text-generation", model=model, tokenizer=tokenizer, device=0,
                        model_kwargs={"torch_dtype":torch.bfloat16})

    def _call(self, prompt, stop=None):
        return self.pipeline(prompt, max_length=9999)[0]["generated_text"]
 
    def _identifying_params(self):
        return {"name_of_model": self.model_name}

    def _llm_type(self):
        return "custom"


llm_predictor = LLMPredictor(llm=CustomLLM())

运行时出现CUDA内存不足错误:

OutOfMemoryError: CUDA out of memory. Tried to allocate 270.00 MiB (GPU 0; 22.03 GiB total capacity; 21.65 GiB 
already allocated; 94.88 MiB free; 21.65 GiB reserved in total by PyTorch) If reserved memory is >> allocated 
memory try setting max_split_size_mb to avoid fragmentation.  See documentation for Memory Management and 
PYTORCH_CUDA_ALLOC_CONF

运行前!nvidia-smi输出:

Thu Apr 20 18:04:00 2023       
+---------------------------------------------------------------------------------------+
| NVIDIA-SMI 530.30.02              Driver Version: 530.30.02    CUDA Version: 12.1     |
|-----------------------------------------+----------------------+----------------------+
| GPU  Name                  Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf            Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                                         |                      |               MIG M. |
|=========================================+======================+======================|
|   0  NVIDIA A10G                     Off| 00000000:00:1E.0 Off |                    0 |
|  0%   23C    P0               52W / 300W|      0MiB / 23028MiB |     18%      Default |
|                                         |                      |                  N/A |
+-----------------------------------------+----------------------+----------------------+
                                                                                         
+---------------------------------------------------------------------------------------+
| Processes:                                                                            |
|  GPU   GI   CI        PID   Type   Process name                            GPU Memory |
|        ID   ID                                                             Usage      |
|=======================================================================================|
|  No running processes found                                                           |
+---------------------------------------------------------------------------------------+
解决方案

针对Vicuna-13B在A10G(22GB显存)上的OOM问题,可通过以下代码调整解决:

  • 启用4-bit量化:借助bitsandbytes库实现模型量化,大幅降低显存占用。
    先安装依赖:

    pip install bitsandbytes accelerate
    

    修改后的完整代码:

    import os
    import torch
    from llama_index import LLMPredictor
    from langchain.llms.base import LLM
    from transformers import pipeline, AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
    
    # 正确设置显存配置环境变量
    os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:512"
    torch.cuda.empty_cache()
    
    # 配置4-bit量化参数
    bnb_config = BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_use_double_quant=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16
    )
      
    class CustomLLM(LLM):
        model_name = "eachadea/vicuna-13b-1.1"
        tokenizer = AutoTokenizer.from_pretrained(model_name)
        # 加载量化后的模型,自动分配设备
        model = AutoModelForCausalLM.from_pretrained(
            model_name,
            quantization_config=bnb_config,
            device_map="auto",
            torch_dtype=torch.bfloat16
        )
    
        # Vicuna是自回归模型,任务类型改为text-generation
        pipeline = pipeline("text-generation", model=model, tokenizer=tokenizer, 
                            device_map="auto",
                            model_kwargs={"torch_dtype":torch.bfloat16})
    
        def _call(self, prompt, stop=None):
            # 限制生成新token数量,避免显存溢出
            outputs = self.pipeline(prompt, max_new_tokens=512, do_sample=True, temperature=0.7)
            return outputs[0]["generated_text"].replace(prompt, "")
    
        def _identifying_params(self):
            return {"name_of_model": self.model_name}
    
        def _llm_type(self):
            return "custom"
    
    
    llm_predictor = LLMPredictor(llm=CustomLLM())
    
  • 关键调整说明:

    • 替换text2text-generation为text-generation,匹配Vicuna模型的任务类型
    • 使用max_new_tokens替代max_length,避免计算冗余上下文长度
    • 加载模型时添加device_map="auto",让框架自动分配模型层到GPU/CPU
    • 代码开头添加torch.cuda.empty_cache()清理显存碎片

内容的提问来源于stack exchange,提问作者Ben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 14:17:23