部署LangChain+Dash问答机器人时CUDA内存不足问题求解
问题
我正在构建基于LangChain的问答机器人,并用Python Dash部署。运行时出现CUDA内存不足错误:
torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 64.00 MiB (GPU 0; 4.00 GiB total capacity; 3.44 GiB already allocated; 0 bytes free; 3.44 GiB reserved in total by PyTorch)
If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF
机器人在CPU上运行正常,但为了提升扩展性想启用CUDA加速。我已经尝试过以下方法但都没解决:
- 设置
PYTORCH_CUDA_ALLOC_CONF为512mb; - 设置
batch_size=1; - 在chain_type的'stuff'和'map_reduce'之间切换。
相关代码如下:
vector_db = Chroma( persist_directory = "", embedding_function = HuggingFaceInstructEmbeddings( model_name = "hkunlp/instructor-xl", model_kwargs = { "device": "cuda" })) llm = AzureOpenAI("",batch_size=1) qa_chain = RetrievalQA.from_chain_type( llm = llm, chain_type = "map_reduce", retriever = vector_db.as_retriever( search_kwargs = { 'k': 1 }), return_source_documents = True)
解决方案
- 换用轻量嵌入模型:
hkunlp/instructor-xl属于大参数模型,4GB显存无法承载。换成显存占用更低的同系列模型hkunlp/instructor-base,或者通用轻量模型sentence-transformers/all-MiniLM-L6-v2,能在保证嵌入效果的前提下大幅降低显存消耗。修改代码示例:embedding_function = HuggingFaceInstructEmbeddings( model_name = "hkunlp/instructor-base", model_kwargs = { "device": "cuda" }) - 启用模型量化:如果必须使用
instructor-xl,可以通过量化压缩模型体积。先安装bitsandbytes库,再在模型参数中开启8位或4位量化:embedding_function = HuggingFaceInstructEmbeddings( model_name = "hkunlp/instructor-xl", model_kwargs = { "device": "cuda", "load_in_8bit": True }) - 手动清理CUDA缓存:在嵌入生成、问答请求处理的关键节点,手动释放无用显存:
可以将这段代码放在每次调用import torch torch.cuda.empty_cache()qa_chain之后,或者嵌入模型初始化完成后。 - 拆分批量任务:如果处理的文档量较大,不要一次性加载所有文档生成嵌入,分批次处理,每完成一批就清理一次缓存。
内容的提问来源于stack exchange,提问作者jvlmtc
相关产品推荐
相关产品推荐

