使用Langchain处理多PDF时摘要异常,重复输出首个PDF结果
解决Chroma向量数据库复用旧PDF向量的问题
问题根源在于你将Chroma的向量数据持久化到了当前目录(.),每次运行代码时,Chroma会自动加载该目录下已存在的向量数据,而非完全基于新PDF生成新向量,导致新PDF的摘要仍使用旧数据。
解决方案:
清理旧的持久化数据
每次处理新PDF前,删除当前目录下的Chroma相关数据文件(默认是chroma.sqlite3、chroma-collections.parquet等),也可以用代码自动清理:import os import shutil # 清理当前目录下的Chroma数据文件 persist_dir = "." for item in os.listdir(persist_dir): item_path = os.path.join(persist_dir, item) if item.startswith("chroma"): if os.path.isfile(item_path): os.remove(item_path) elif os.path.isdir(item_path): shutil.rmtree(item_path)将这段代码放在创建
vectordb之前,确保每次运行都从空状态开始。使用内存存储(不持久化)
若不需要长期保存向量数据,去掉persist_directory参数,Chroma会将向量存在内存中,程序结束后自动销毁:vectordb = Chroma.from_documents(pages, embedding=embeddings) # 无需调用vectordb.persist()为不同PDF分配独立持久化目录
给每个PDF单独指定存储目录,避免数据混淆:# 根据PDF文件名生成专属目录 pdf_name = os.path.basename(pdf_path).split(".")[0] persist_dir = f"./chroma_storage_{pdf_name}" vectordb = Chroma.from_documents(pages, embedding=embeddings, persist_directory=persist_dir) vectordb.persist()
修复后的完整代码示例(自动清理旧数据版本):
from langchain.document_loaders import PyPDFLoader from langchain.embeddings import OpenAIEmbeddings from langchain.vectorstores import Chroma from langchain.chains import ChatVectorDBChain from langchain.llms import OpenAI import os import shutil os.environ["OPENAI_API_KEY"] = "my_API_KEY" pdf_path = "new_file_path" loader = PyPDFLoader(pdf_path) pages = loader.load_and_split() print(pages[1].page_content) # 清理旧的Chroma持久化数据 persist_dir = "." for item in os.listdir(persist_dir): item_path = os.path.join(persist_dir, item) if item.startswith("chroma"): if os.path.isfile(item_path): os.remove(item_path) elif os.path.isdir(item_path): shutil.rmtree(item_path) embeddings = OpenAIEmbeddings() vectordb = Chroma.from_documents(pages, embedding=embeddings, persist_directory=persist_dir) vectordb.persist() pdf_qa = ChatVectorDBChain.from_llm(OpenAI(temperature=0.9, model_name="gpt-3.5-turbo"), vectordb, return_source_documents=True) query = "Write a summary of the text." result = pdf_qa({"question": query, "chat_history": ""}) print(result["answer"])
内容的提问来源于stack exchange,提问作者rna_2090
相关产品推荐
相关产品推荐

