向Azure Cognitive Search索引加载含嵌入向量数据时遇错误求助
问题:向Azure Cognitive Search索引加载含嵌入向量数据失败
问题场景
尝试将包含嵌入向量的Pandas DataFrame数据加载到Azure Cognitive Search索引,添加嵌入字段后执行批量上传和直接上传操作均报错。
批量上传代码及错误
批量上传代码
input_data = df.to_json() # DF为包含嵌入字段的Pandas DataFrame # 使用SearchIndexingBufferedSender批量上传优化索引 with SearchIndexingBufferedSender( endpoint=service_endpoint, index_name=index_name, credential=credential, ) as batch_client: # 添加所有文档的上传操作 batch_client.upload_documents(documents=input_data) print(f"已上传共{len(input_data)}份文档")
错误信息
File /packages/azure/search/documents/_search_indexing_buffered_sender.py:322, in SearchIndexingBufferedSender._retry_action(self, action) 320 self._callback_fail(action) 321 return --> 322 key = action.additional_properties.get(self._index_key) 323 counter = self._retry_counter.get(key) 324 if not counter: 325 # 首次失败 AttributeError: 'str' object has no attribute 'get'
直接上传代码及错误
直接上传代码
search_client = SearchClient(endpoint=service_endpoint, index_name=index_name, credential=credential) result = search_client.upload_documents(input_data, timeout = 50)
错误信息
File /packages/azure/search/documents/_generated/operations/_documents_operations.py:1251, in DocumentsOperations.index(self, batch, request_options, **kwargs) 1249 map_error(status_code=response.status_code, response=response, error_map=error_map) 1250 error = self._deserialize.failsafe_deserialize(_models.SearchError, pipeline_response) --> 1251 raise HttpResponseError(response=response, model=error) 1253 if response.status_code == 200: 1254 deserialized = self._deserialize("IndexDocumentsResult", pipeline_response) HttpResponseError: () 请求无效。详细信息:发现预期类型'search.documentFields[Nullable=False]'存在空值。预期类型'search.documentFields[Nullable=False]'不允许空值。 Code: Message: 请求无效。详细信息:发现预期类型'search.documentFields[Nullable=False]'存在空值。预期类型'search.documentFields[Nullable=False]'不允许空值。
尝试过的格式调整
尝试两种JSON格式转换,均未解决问题:
input_data = df.to_json() input_data = df.to_json(orient="records")
索引定义
index_client = SearchIndexClient( endpoint=service_endpoint, credential=credential) fields = [ SimpleField(name="Id", type=SearchFieldDataType.String, key=True, sortable=True, filterable=True, facetable=True), SearchableField(name="Field1", type=SearchFieldDataType.String), SearchableField(name="Field2", type=SearchFieldDataType.String, filterable=True), SearchableField(name="Field3", type=SearchFieldDataType.String, filterable=True), SearchableField(name="Field4", type=SearchFieldDataType.String, filterable=True), SearchableField(name="Field5", type=SearchFieldDataType.String, filterable=True), SearchField(name="Field4_vec", type=SearchFieldDataType.Collection(SearchFieldDataType.Single), searchable=True, vector_search_dimensions=384, vector_search_profile="myHnswProfile"), SearchField(name="Field5_vec", type=SearchFieldDataType.Collection(SearchFieldDataType.Single), searchable=True, vector_search_dimensions=384, vector_search_profile="myHnswProfile") ] # 配置向量搜索设置 vector_search = VectorSearch( algorithms=[ HnswVectorSearchAlgorithmConfiguration( name="myHnsw", kind=VectorSearchAlgorithmKind.HNSW, parameters=HnswParameters( m=4, ef_construction=400, ef_search=500, metric="cosine" ) ), ExhaustiveKnnVectorSearchAlgorithmConfiguration( name="myExhaustiveKnn", kind=VectorSearchAlgorithmKind.EXHAUSTIVE_KNN, parameters=ExhaustiveKnnParameters( metric="cosine" ) ) ], profiles=[ VectorSearchProfile( name="myHnswProfile", algorithm="myHnsw", ), VectorSearchProfile( name="myExhaustiveKnnProfile", algorithm="myExhaustiveKnn", ) ] ) # 创建搜索索引 index = SearchIndex(name=index_name, fields=fields, vector_search=vector_search) result = index_client.create_or_update_index(index) print(f' {result.name} 创建成功')
DataFrame字段结构
DataFrame字段与索引完全匹配:
- Id (字符串)
- Field1 (字符串)
- Field2 (字符串)
- Field3 (字符串)
- Field4 (字符串)
- Field5 (字符串)
- Field4_vec (384维浮点型向量,格式如
[-0.01168345008045435, -0.0396871380507946, ...]) - Field5_vec (384维浮点型向量,格式同Field4_vec)
解决方案
核心问题
Azure Cognitive Search的upload_documents方法接受Python字典列表作为参数,而非JSON字符串。之前用df.to_json()得到的是字符串,导致批量上传时触发类型错误,直接上传时因格式解析问题误判为空值。
修正后的代码
1. 批量上传方式
# 将DataFrame转为字典列表,而非JSON字符串 input_data = df.to_dict(orient="records") with SearchIndexingBufferedSender( endpoint=service_endpoint, index_name=index_name, credential=credential, ) as batch_client: batch_client.upload_documents(documents=input_data) print(f"已上传共{len(input_data)}份文档")
2. 直接上传方式
search_client = SearchClient(endpoint=service_endpoint, index_name=index_name, credential=credential) # 转为字典列表 input_data = df.to_dict(orient="records") result = search_client.upload_documents(input_data, timeout=50) # 验证上传结果 for item in result: if not item.succeeded: print(f"文档{item.key}上传失败: {item.error_message}")
额外检查点
- 确认向量字段
Field4_vec和Field5_vec是Python列表类型,而非字符串形式的数组。如果是字符串,需先解析:import ast df['Field4_vec'] = df['Field4_vec'].apply(ast.literal_eval) df['Field5_vec'] = df['Field5_vec'].apply(ast.literal_eval) - 验证所有字段无空值,主键
Id唯一且非空:print(df.isnull().sum()) print(df['Id'].duplicated().sum())
内容的提问来源于stack exchange,提问作者Irina
相关产品推荐
相关产品推荐

