如何捕获langchain中Chroma.from_documents的重复ID错误
处理Chroma重复ID报错并避免重复添加文档
直接捕获错误忽略重复项
你可以通过Python的try-except语句直接捕获chromadb.errors.IDAlreadyExistsError异常,遇到重复ID时跳过对应文档或执行自定义逻辑。
首先导入对应的异常类:
from chromadb.errors import IDAlreadyExistsError
然后用try-except包裹添加逻辑:
from langchain.vectorstores import Chroma from langchain.embeddings import YourEmbeddingModel # 替换为你实际使用的嵌入模型 from chromadb.errors import IDAlreadyExistsError # 假设docs、embeddings、ids已定义 try: db = Chroma.from_documents(docs, embeddings, ids=ids, persist_directory='db') db.persist() except IDAlreadyExistsError as e: # 可添加自定义处理逻辑,比如打印提示、记录日志 print(f"跳过重复ID: {e}") # 如需保留已添加文档,可加载已存在的数据库 db = Chroma(persist_directory='db', embedding_function=embeddings)
提前检查ID避免报错
如果想在添加前就过滤掉已存在的ID,可先加载现有Chroma数据库,获取所有已存在的ID,再过滤待添加的文档和ID列表:
from langchain.vectorstores import Chroma from langchain.embeddings import YourEmbeddingModel # 加载已有的数据库 existing_db = Chroma(persist_directory='db', embedding_function=embeddings) existing_ids = set(existing_db.get()['ids']) # 过滤重复项 new_docs = [] new_ids = [] for doc, doc_id in zip(docs, ids): if doc_id not in existing_ids: new_docs.append(doc) new_ids.append(doc_id) # 仅添加新文档 if new_docs: db = Chroma.from_documents(new_docs, embeddings, ids=new_ids, persist_directory='db') db.persist() else: print("没有新文档需要添加")
逐文档添加跳过重复项
Chroma.from_documents是批量操作,单个重复ID会导致整个批量失败。若要批量中跳过重复项单独添加,可循环遍历每个文档和ID:
from langchain.vectorstores import Chroma from langchain.embeddings import YourEmbeddingModel from chromadb.errors import IDAlreadyExistsError # 初始化或加载数据库 db = Chroma(persist_directory='db', embedding_function=embeddings) for doc, doc_id in zip(docs, ids): try: db.add_documents([doc], ids=[doc_id]) except IDAlreadyExistsError: print(f"ID {doc_id} 已存在,跳过") db.persist()
内容的提问来源于stack exchange,提问作者Suibhne
相关产品推荐
相关产品推荐

