You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Langchain开发时内存持续上涨问题求助

解决Flask中Langchain向量数据库内存持续上涨的问题

核心问题分析

你的代码每次请求都会新建RecursiveCharacterTextSplitter、OpenAIEmbeddings实例,构建完整FAISS索引后直接保存,但这些对象在请求结束后可能未被Python垃圾回收(GC)及时清理。加上Flask的运行机制(比如多线程模式下线程残留对象),导致内存不断累积。更换向量库无效,说明问题不在向量库本身,而是每次请求重复创建资源且未合理释放的通用问题。

具体解决方案

1. 复用Embeddings实例

将OpenAIEmbeddings实例全局初始化,避免每次请求重复创建(Embeddings初始化会加载相关资源,重复创建会造成内存冗余):

# 全局初始化,仅创建一次
embeddings = OpenAIEmbeddings()

def upload_data():
    text = request.get_json().get('text')
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=150)
    docs = text_splitter.split_text(text)
    
    # 复用全局embeddings实例
    faiss_db = FAISS.from_texts(docs, embeddings)
    faiss_db.save_local("faiss_index")
    
    # 显式删除局部变量,辅助GC回收
    del faiss_db, docs
    return "Success"

2. 手动触发垃圾回收

在请求结束后强制调用GC,清理未被引用的对象:

import gc

def upload_data():
    text = request.get_json().get('text')
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=150)
    docs = text_splitter.split_text(text)
    
    faiss_db = FAISS.from_texts(docs, embeddings)
    faiss_db.save_local("faiss_index")
    
    # 清理局部变量
    del faiss_db, docs, text_splitter
    # 强制触发GC
    gc.collect()
    return "Success"

3. 增量更新索引(可选优化)

如果是增量上传场景,不要每次重建完整索引,而是加载已有索引后添加新文档,大幅减少内存开销:

def upload_data():
    text = request.get_json().get('text')
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=150)
    docs = text_splitter.split_text(text)
    
    try:
        # 加载已存在的索引
        faiss_db = FAISS.load_local("faiss_index", embeddings, allow_dangerous_deserialization=True)
        # 添加新文档
        faiss_db.add_texts(docs)
    except FileNotFoundError:
        # 首次创建索引
        faiss_db = FAISS.from_texts(docs, embeddings)
    
    faiss_db.save_local("faiss_index")
    del faiss_db, docs
    gc.collect()
    return "Success"

4. 调整Flask运行配置

  • 关闭debug模式:debug模式会启动额外监控进程,加剧内存泄漏,生产环境必须禁用:
if __name__ == '__main__':
    app.run(debug=False)
  • 使用生产级服务器:比如Gunicorn,配合合理的worker数量,避免单进程内存无限累积。

5. 禁用Embeddings缓存

OpenAIEmbeddings默认会缓存生成的向量,长期运行会占用大量内存,不需要缓存时可直接禁用:

embeddings = OpenAIEmbeddings(cache=False)

内容的提问来源于stack exchange,提问作者System Life

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 14:13:23