Django中LangChain+Chroma PersistentClient存储PDF至ChromaDB失败排查
问题
在Django项目中尝试用LangChain将PDF文档嵌入ChromaDB向量数据库,采用PersistentClient实现,但Chroma未成功保存文档。执行代码后打印集合内文档结果为空数组[],已确认documents、strip_user_email(user.email)、embedding和client均为有效变量。
相关代码:
from rest_framework.response import Response from rest_framework import viewsets from langchain.chat_models import ChatOpenAI import chromadb from ..functions.load_new_pdf import load_new_pdf from ..functions.store_docs_vector import store_embeds import sys from ..models import Documents from .BaseView import get_user, strip_user_email from ..functions.llm import chosen_llm from langchain_community.vectorstores import Chroma from langchain.embeddings.openai import OpenAIEmbeddings from dotenv import load_dotenv import sys import os load_dotenv() OPENAI_API_KEY = os.getenv('OPENAI_API_KEY') embedding = OpenAIEmbeddings() llm = chosen_llm() def uploadDocument(self, request): user = get_user(request) base64_data = request.data.get('file') # Create a record of the document in the SQLite database. document = Documents.objects.create(name=request.data.get( "name"), user=user, content=base64_data) document.save() # Convert the pdf into a big string. cleaned_text, documents = load_new_pdf(document) # Initiliaze the persistent client. client = chromadb.PersistentClient(path="../chroma") # Strip the user email from the '@' and '.' characters and get or create a new collection for it collection_name = strip_user_email(user.email) client.get_or_create_collection(collection_name) # Embed the documents into the database Chroma.from_documents( documents=documents, embedding=embedding, client=client) # Retrieve the collection from the database chroma_db = Chroma(collection_name=collection_name, embedding_function=embedding, client=client) print(chroma_db.get()["documents"], file=sys.stderr) # This is now [], why??
问题原因与解决方法
未指定目标集合名称
调用Chroma.from_documents时没有传入collection_name参数,LangChain会默认创建一个名为langchain的集合,而非你通过client.get_or_create_collection创建的用户专属集合,导致后续查询用户集合时为空。
解决:在Chroma.from_documents中添加collection_name参数:Chroma.from_documents( documents=documents, embedding=embedding, client=client, collection_name=collection_name )相对路径可能导致存储位置错误
Django项目中../chroma的相对路径可能与预期不符,每次运行时创建新的Chroma实例,无法读取之前的存储数据。
解决:改用基于项目根目录的绝对路径:import os from django.conf import settings chroma_path = os.path.join(settings.BASE_DIR, "chroma") client = chromadb.PersistentClient(path=chroma_path)需确认
documents的格式正确性
即使documents是有效变量,也要确保它是LangChain的Document对象列表,而非纯文本字符串。如果load_new_pdf返回的是纯文本,需要先转换:from langchain_core.documents import Document # 假设cleaned_text是纯文本,转换为Document列表 documents = [Document(page_content=cleaned_text)]
内容的提问来源于stack exchange,提问作者Yassine Mabrouk
相关产品推荐
相关产品推荐

