You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Django中LangChain+Chroma PersistentClient存储PDF至ChromaDB失败排查

问题

在Django项目中尝试用LangChain将PDF文档嵌入ChromaDB向量数据库,采用PersistentClient实现,但Chroma未成功保存文档。执行代码后打印集合内文档结果为空数组[],已确认documents、strip_user_email(user.email)、embedding和client均为有效变量。

相关代码:

from rest_framework.response import Response
from rest_framework import viewsets
from langchain.chat_models import ChatOpenAI
import chromadb

from ..functions.load_new_pdf import load_new_pdf
from ..functions.store_docs_vector import store_embeds
import sys

from ..models import Documents

from .BaseView import get_user, strip_user_email

from ..functions.llm import chosen_llm

from langchain_community.vectorstores import Chroma
from langchain.embeddings.openai import OpenAIEmbeddings
from dotenv import load_dotenv
import sys
import os
load_dotenv()

OPENAI_API_KEY = os.getenv('OPENAI_API_KEY')


embedding = OpenAIEmbeddings()


llm = chosen_llm()


def uploadDocument(self, request):

    user = get_user(request)
    base64_data = request.data.get('file')

    # Create a record of the document in the SQLite database.
    document = Documents.objects.create(name=request.data.get(
        "name"), user=user, content=base64_data)
    document.save()

    # Convert the pdf into a big string.
    cleaned_text, documents = load_new_pdf(document)

    # Initiliaze the persistent client. 
    client = chromadb.PersistentClient(path="../chroma")

    # Strip the user email from the '@' and '.' characters and get or create a new collection for it
    collection_name = strip_user_email(user.email)
    client.get_or_create_collection(collection_name)

    # Embed the documents into the database
    Chroma.from_documents(
        documents=documents, embedding=embedding,
        client=client)

    # Retrieve the collection from the database
    chroma_db = Chroma(collection_name=collection_name, embedding_function=embedding, client=client)

    print(chroma_db.get()["documents"], file=sys.stderr) # This is now [], why??

问题原因与解决方法

  • 未指定目标集合名称
    调用Chroma.from_documents时没有传入collection_name参数,LangChain会默认创建一个名为langchain的集合,而非你通过client.get_or_create_collection创建的用户专属集合,导致后续查询用户集合时为空。
    解决:在Chroma.from_documents中添加collection_name参数:

    Chroma.from_documents(
        documents=documents, embedding=embedding,
        client=client, collection_name=collection_name
    )
    
  • 相对路径可能导致存储位置错误
    Django项目中../chroma的相对路径可能与预期不符,每次运行时创建新的Chroma实例,无法读取之前的存储数据。
    解决:改用基于项目根目录的绝对路径:

    import os
    from django.conf import settings
    
    chroma_path = os.path.join(settings.BASE_DIR, "chroma")
    client = chromadb.PersistentClient(path=chroma_path)
    
  • 需确认documents的格式正确性
    即使documents是有效变量,也要确保它是LangChain的Document对象列表,而非纯文本字符串。如果load_new_pdf返回的是纯文本,需要先转换:

    from langchain_core.documents import Document
    
    # 假设cleaned_text是纯文本,转换为Document列表
    documents = [Document(page_content=cleaned_text)]
    

内容的提问来源于stack exchange,提问作者Yassine Mabrouk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 12:03:16