You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

向Azure Cognitive Search索引加载含嵌入向量数据时遇错误求助

问题:向Azure Cognitive Search索引加载含嵌入向量数据失败

问题场景

尝试将包含嵌入向量的Pandas DataFrame数据加载到Azure Cognitive Search索引,添加嵌入字段后执行批量上传和直接上传操作均报错。

批量上传代码及错误

批量上传代码

input_data = df.to_json() # DF为包含嵌入字段的Pandas DataFrame

# 使用SearchIndexingBufferedSender批量上传优化索引
with SearchIndexingBufferedSender(  
    endpoint=service_endpoint,  
    index_name=index_name,  
    credential=credential,  
) as batch_client:  
    # 添加所有文档的上传操作
    batch_client.upload_documents(documents=input_data)  
print(f"已上传共{len(input_data)}份文档")

错误信息

File /packages/azure/search/documents/_search_indexing_buffered_sender.py:322, in SearchIndexingBufferedSender._retry_action(self, action)
    320     self._callback_fail(action)
    321     return
--> 322 key = action.additional_properties.get(self._index_key)
    323 counter = self._retry_counter.get(key)
    324 if not counter:
    325     # 首次失败

AttributeError: 'str' object has no attribute 'get'

直接上传代码及错误

直接上传代码

search_client = SearchClient(endpoint=service_endpoint, index_name=index_name, credential=credential)
result = search_client.upload_documents(input_data, timeout = 50)

错误信息

File /packages/azure/search/documents/_generated/operations/_documents_operations.py:1251, in DocumentsOperations.index(self, batch, request_options, **kwargs)
   1249     map_error(status_code=response.status_code, response=response, error_map=error_map)
   1250     error = self._deserialize.failsafe_deserialize(_models.SearchError, pipeline_response)
--> 1251     raise HttpResponseError(response=response, model=error)
   1253 if response.status_code == 200:
   1254     deserialized = self._deserialize("IndexDocumentsResult", pipeline_response)

HttpResponseError: () 请求无效。详细信息:发现预期类型'search.documentFields[Nullable=False]'存在空值。预期类型'search.documentFields[Nullable=False]'不允许空值。
Code: 
Message: 请求无效。详细信息:发现预期类型'search.documentFields[Nullable=False]'存在空值。预期类型'search.documentFields[Nullable=False]'不允许空值。

尝试过的格式调整

尝试两种JSON格式转换,均未解决问题:

input_data = df.to_json()
input_data = df.to_json(orient="records")

索引定义

index_client = SearchIndexClient(
    endpoint=service_endpoint, credential=credential)

fields = [
    SimpleField(name="Id", type=SearchFieldDataType.String, key=True, sortable=True, filterable=True, facetable=True),
    SearchableField(name="Field1", type=SearchFieldDataType.String),
    SearchableField(name="Field2", type=SearchFieldDataType.String, filterable=True),
    SearchableField(name="Field3", type=SearchFieldDataType.String, filterable=True),
    SearchableField(name="Field4", type=SearchFieldDataType.String, filterable=True),
    SearchableField(name="Field5", type=SearchFieldDataType.String, filterable=True),
    SearchField(name="Field4_vec", type=SearchFieldDataType.Collection(SearchFieldDataType.Single),
                searchable=True, vector_search_dimensions=384, vector_search_profile="myHnswProfile"),
    SearchField(name="Field5_vec", type=SearchFieldDataType.Collection(SearchFieldDataType.Single),
                searchable=True, vector_search_dimensions=384, vector_search_profile="myHnswProfile")
]

# 配置向量搜索设置
vector_search = VectorSearch(
    algorithms=[
        HnswVectorSearchAlgorithmConfiguration(
            name="myHnsw",
            kind=VectorSearchAlgorithmKind.HNSW,
            parameters=HnswParameters(
                m=4,
                ef_construction=400,
                ef_search=500,
                metric="cosine"
            )
        ),
        ExhaustiveKnnVectorSearchAlgorithmConfiguration(
            name="myExhaustiveKnn",
            kind=VectorSearchAlgorithmKind.EXHAUSTIVE_KNN,
            parameters=ExhaustiveKnnParameters(
                metric="cosine"
            )
        )
    ],
    profiles=[
        VectorSearchProfile(
            name="myHnswProfile",
            algorithm="myHnsw",
        ),
        VectorSearchProfile(
            name="myExhaustiveKnnProfile",
            algorithm="myExhaustiveKnn",
        )
    ]
)

# 创建搜索索引
index = SearchIndex(name=index_name, fields=fields,
                    vector_search=vector_search)
result = index_client.create_or_update_index(index)
print(f' {result.name} 创建成功')

DataFrame字段结构

DataFrame字段与索引完全匹配:

  • Id (字符串)
  • Field1 (字符串)
  • Field2 (字符串)
  • Field3 (字符串)
  • Field4 (字符串)
  • Field5 (字符串)
  • Field4_vec (384维浮点型向量,格式如[-0.01168345008045435, -0.0396871380507946, ...])
  • Field5_vec (384维浮点型向量,格式同Field4_vec)

解决方案

核心问题

Azure Cognitive Search的upload_documents方法接受Python字典列表作为参数,而非JSON字符串。之前用df.to_json()得到的是字符串,导致批量上传时触发类型错误,直接上传时因格式解析问题误判为空值。

修正后的代码

1. 批量上传方式

# 将DataFrame转为字典列表,而非JSON字符串
input_data = df.to_dict(orient="records")

with SearchIndexingBufferedSender(  
    endpoint=service_endpoint,  
    index_name=index_name,  
    credential=credential,  
) as batch_client:  
    batch_client.upload_documents(documents=input_data)  
print(f"已上传共{len(input_data)}份文档")

2. 直接上传方式

search_client = SearchClient(endpoint=service_endpoint, index_name=index_name, credential=credential)
# 转为字典列表
input_data = df.to_dict(orient="records")
result = search_client.upload_documents(input_data, timeout=50)

# 验证上传结果
for item in result:
    if not item.succeeded:
        print(f"文档{item.key}上传失败: {item.error_message}")

额外检查点

  • 确认向量字段Field4_vec和Field5_vec是Python列表类型,而非字符串形式的数组。如果是字符串,需先解析:
    import ast
    df['Field4_vec'] = df['Field4_vec'].apply(ast.literal_eval)
    df['Field5_vec'] = df['Field5_vec'].apply(ast.literal_eval)
    
  • 验证所有字段无空值,主键Id唯一且非空:
    print(df.isnull().sum())
    print(df['Id'].duplicated().sum())
    

内容的提问来源于stack exchange,提问作者Irina

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 18:54:52