使用Milvus与Haystack时出现重复集合创建错误
Haystack + Milvus Lite 写入时重复字段错误的解决方案
问题背景
在Haystack中使用Milvus作为文档存储,执行索引写入时触发重复字段错误,调试确认是重复创建默认集合导致。
相关代码
MilvusDocumentStore 连接代码
@lru_cache def get_vector_db(): # Get document store from database return MilvusDocumentStore( connection_args={ "uri": get_settings().milvus_db_path }, # Milvus Lite drop_old=True )
索引管道定义代码
file_type_router = FileTypeRouter( mime_types=[ "text/plain" ] ) # Converter plain text files to Document objects text_converter = TextFileToDocument() # Join Documents coming from different branches of a pipeline document_joiner = DocumentJoiner() # Clean the text of the documents document_cleaner = DocumentCleaner() # Split the documents into smaller documents document_splitter = DocumentSplitter(split_by="sentence", split_length=2) # Create embeddings from the Documents document_embedder = SentenceTransformersDocumentEmbedder( model="sentence-transformers/all-MiniLM-L6-v2" ) # Write the documents to the DocumentStore document_writer = DocumentWriter(document_store, policy=DuplicatePolicy.NONE) # Build the Indexing pipeline preprocessing_pipeline = Pipeline() preprocessing_pipeline.add_component( name="file_type_router", instance=file_type_router ) preprocessing_pipeline.add_component(name="text_converter", instance=text_converter) preprocessing_pipeline.add_component( name="document_joiner", instance=document_joiner ) preprocessing_pipeline.add_component( name="document_cleaner", instance=document_cleaner ) preprocessing_pipeline.add_component( name="document_splitter", instance=document_splitter ) preprocessing_pipeline.add_component( name="document_embedder", instance=document_embedder ) preprocessing_pipeline.add_component( name="document_writer", instance=document_writer ) # Connect components preprocessing_pipeline.connect( "file_type_router.plain/text", "text_converter.sources" ) preprocessing_pipeline.connect("text_converter", "document_joiner") preprocessing_pipeline.connect("document_joiner", "document_cleaner") preprocessing_pipeline.connect("document_cleaner", "document_splitter") preprocessing_pipeline.connect("document_splitter", "document_embedder") preprocessing_pipeline.connect("document_embedder", "duplicate_checker") preprocessing_pipeline.connect( "duplicate_checker.documents_to_index", "document_writer.documents" )
错误信息
Failed to create collection: HaystackCollection error: <MilvusException: (code=2000, message=Assert "!name_ids_.count(field_name)" at /Users/zilliz/milvus-lite/thirdparty/milvus/internal/core/src/common/Schema.h:172 => duplicated field name: segcore error)> ERROR: Exception in ASGI application
解决方法
1. 修改drop_old参数
当前MilvusDocumentStore设置drop_old=True,每次初始化都会删除旧集合并重建,若集合已存在且字段结构冲突就会触发错误。建议改为drop_old=False,仅在首次创建或需要重置集合时设为True:
@lru_cache def get_vector_db(): return MilvusDocumentStore( connection_args={ "uri": get_settings().milvus_db_path }, drop_old=False # 禁用自动删除旧集合 )
2. 修复管道组件缺失问题
索引管道中连接了duplicate_checker但未添加该组件,导致逻辑异常。需补充组件定义:
# 添加DuplicateChecker组件 duplicate_checker = DuplicateChecker(document_store=document_store) preprocessing_pipeline.add_component(name="duplicate_checker", instance=duplicate_checker)
若不需要重复检查逻辑,可直接修改连接关系:
preprocessing_pipeline.connect("document_embedder", "document_writer.documents")
3. 手动清理旧集合(可选)
如果需要彻底重置集合,可在初始化时手动删除旧集合再重建:
@lru_cache def get_vector_db(): document_store = MilvusDocumentStore( connection_args={ "uri": get_settings().milvus_db_path }, drop_old=False ) # 手动删除默认集合 if document_store.collection_exists("HaystackCollection"): document_store.delete_collection("HaystackCollection") # 重新创建集合 document_store.create_collection() return document_store
内容的提问来源于stack exchange,提问作者cksrc
相关产品推荐
相关产品推荐

