You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Celery任务中提升ChromaDB文档写入速度?

解决ChromaDB在Celery任务中重复解压模型导致写入缓慢的问题

核心原因

Celery的每个worker进程(或单次任务执行)都会重新初始化Chroma客户端,而Chroma默认的嵌入模型all-MiniLM-L6-v2会重复触发下载、解压流程,没有在进程间复用已加载的模型资源,导致单条写入的额外开销过大。

具体解决方案

1. 预下载并固定模型缓存

手动完成模型的下载与解压,确保所有Celery worker能直接复用已缓存的模型文件:

  • 找到Chroma默认缓存路径:~/.cache/chroma/onnx_models/
  • 手动下载all-MiniLM-L6-v2的模型包并解压到对应目录,验证/root/.cache/chroma/onnx_models/all-MiniLM-L6-v2/下存在已解压的模型文件(如model.onnx等)
  • 后续任务执行时,Chroma会直接读取本地缓存,跳过下载解压步骤

2. 在Celery Worker启动时预初始化Chroma资源

利用Celery的worker启动钩子,在worker进程启动时一次性初始化Chroma客户端和嵌入模型,所有任务复用同一实例:

from celery import Celery
import chromadb
from chromadb.utils import embedding_functions

app = Celery('tasks', broker='your_broker_address')

# 全局变量存储预初始化的资源
chroma_client = None
embedding_fn = None

@app.on_after_configure.connect
def init_chroma_resources(sender, **kwargs):
    global chroma_client, embedding_fn
    # 初始化嵌入函数,复用本地缓存模型
    embedding_fn = embedding_functions.SentenceTransformerEmbeddingFunction(
        model_name="all-MiniLM-L6-v2"
    )
    chroma_client = chromadb.PersistentClient(path="./local_chroma_storage")

@app.task
def add_single_document(content):
    global chroma_client, embedding_fn
    collection = chroma_client.get_or_create_collection(
        name="target_collection", 
        embedding_function=embedding_fn
    )
    collection.add(documents=[content], ids=[f"doc_{hash(content)}"])

每个worker进程仅初始化一次模型,所有任务共享该资源。

3. 批量写入替代单条任务

避免单条文档触发一个Celery任务,积累一批文档后批量写入,减少任务启动和模型初始化的重复开销:

@app.task
def batch_add_documents(documents_list):
    global chroma_client, embedding_fn
    collection = chroma_client.get_or_create_collection(
        name="target_collection", 
        embedding_function=embedding_fn
    )
    doc_ids = [f"doc_{idx}" for idx in range(len(documents_list))]
    collection.add(documents=documents_list, ids=doc_ids)

4. 优化Celery Worker配置

  • 减少worker进程数:比如设置celery -A tasks worker --concurrency=4(根据服务器CPU核心数调整),避免过多进程重复加载模型占用资源
  • 使用线程池模式:添加--pool=threads参数,线程共享进程内的模型资源,彻底避免多进程重复初始化模型的问题

5. 直接指定本地模型路径

如果已拥有本地模型文件,初始化嵌入函数时直接指定路径,彻底跳过下载解压流程:

embedding_fn = embedding_functions.SentenceTransformerEmbeddingFunction(
    model_path="/root/.cache/chroma/onnx_models/all-MiniLM-L6-v2/"
)

内容的提问来源于stack exchange,提问作者Yevhen Vitruk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 06:34:52