向Qdrant添加向量时触发422验证错误,求问题排查方案
尝试向Qdrant添加向量时出现错误,错误信息如下
UnexpectedResponse Traceback (most recent call last)<ipython-input-36-42a89db32382> in ----> 3 add_vectors(embeddings, payload) 6 frames /usr/local/lib/python3.10/dist-packages/qdrant_client/http/api_client.py in send(self, request, type_) 95 except ValidationError as e: 96 raise ResponseHandlingException(e) ---> 97 raise UnexpectedResponse.for_response(response) 98 99 def send_inner(self, request: Request) -> Response: UnexpectedResponse: Unexpected Response: 422 (Unprocessable Entity) Raw response content: b'{"status":{"error":"Validation error in path parameters: [name: value \\"status=<CollectionStatus.GREEN: \\'green\\'> optimizer_status=<OptimizersStatusOneOf.OK: \\'ok\\'> vectors_count=0 indexed_vectors_co ..."}}'
以下是我的代码:
raw_text = "some long text...." def get_chunks(raw_text): text_splitter = CharacterTextSplitter( separator="\n", chunk_size=100, chunk_overlap=50, length_function=len) chunks = text_splitter.split_text(raw_text) return chunks ============ def get_embeddings(chunks, embedding_model_name="text-embedding-ada-002"): points = [] client = OpenAI(api_key=os.environ['OPENAI_API_KEY']) embeddings = [] for chunk in chunks: embeddings.append(client.embeddings.create( input=chunk, model=embedding_model_name).data[0].embedding) return embeddings =========================================================== def add_vectors(vectors, payload): client = qdrant_client.QdrantClient( os.getenv("QDRANT_HOST"), api_key=os.getenv("QDRANT_API_KEY") ) collection = client.get_collection(os.getenv("QDRANT_COLLECTION")) # Create a list of PointStruct objects points = [ models.PointStruct( id=str(i), # Assign unique IDs to points payload=payload, vector=vector ) for i, vector in enumerate(vectors) ] # Insert the points into the vector store client.upsert( collection_name=collection, # Replace with your collection name points=points )
我的调用逻辑如下:
chunks = get_chunks(raw_text) embeddings = get_embeddings(chunks) payload = {"user": "gxxxx"} add_vectors(embeddings, payload)
执行上述调用时触发了上述错误,我已尝试网络上多种解决方案但仍未解决,请问问题出在哪里?
错误原因分析
- collection_name参数类型错误
client.get_collection()返回的是CollectionInfo对象,而client.upsert()的collection_name参数需要字符串类型的集合名称。直接传入对象会导致Qdrant无法解析路径参数,触发422验证错误。 - 代码缩进错误
client.upsert()代码块缩进错误,位于add_vectors函数外部,导致函数执行时不会触发插入操作,且变量作用域存在问题。 - payload复用问题(非直接错误,但需优化)
所有向量点共用同一个payload,未关联对应的文本chunk,后续检索无法获取原始文本内容。
修正后的代码
修正add_vectors函数
def add_vectors(vectors, chunks, base_payload): client = qdrant_client.QdrantClient( os.getenv("QDRANT_HOST"), api_key=os.getenv("QDRANT_API_KEY") ) collection_name = os.getenv("QDRANT_COLLECTION") # 为每个向量点生成包含chunk内容的payload points = [ models.PointStruct( id=str(i), payload={**base_payload, "text": chunk}, # 合并基础payload与当前chunk文本 vector=vector ) for i, (vector, chunk) in enumerate(zip(vectors, chunks)) ] # 执行插入操作,使用字符串类型的集合名称 client.upsert( collection_name=collection_name, points=points )
修正调用逻辑
chunks = get_chunks(raw_text) embeddings = get_embeddings(chunks) base_payload = {"user": "gxxxx"} add_vectors(embeddings, chunks, base_payload)
修正get_chunks函数的缩进错误
def get_chunks(raw_text): text_splitter = CharacterTextSplitter( separator="\n", chunk_size=100, chunk_overlap=50, length_function=len ) chunks = text_splitter.split_text(raw_text) return chunks
关键修正点说明
- 直接使用环境变量中的集合名称字符串,替代
get_collection()返回的对象。 - 调整
client.upsert()的缩进,确保它在add_vectors函数内部执行。 - 将文本chunk传入函数,为每个向量点添加对应的文本内容到payload,提升检索实用性。
- 修复
get_chunks函数的缩进问题,避免语法错误。
内容的提问来源于stack exchange,提问作者Gautam Mukherjee
相关产品推荐
相关产品推荐

