You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure AI Search上传含embedding vector的JSON文档遇类型转换错误

问题解决:Azure AI Search上传向量文档类型不匹配错误

错误根源

错误提示Cannot convert the literal '0.0040753875' to the expected type 'Edm.String',本质是索引架构字段类型定义和上传文档的字段值类型不匹配:要么是索引里把向量字段定义成了字符串类型,要么是上传的向量值是字符串格式而非数值数组。

修复步骤

1. 修正索引架构中的向量字段定义

Azure AI Search的向量字段必须定义为Collection(Edm.Single)类型(浮点数组),绝对不能设为Edm.String。以下是正确的Python库创建索引示例:

from azure.core.credentials import AzureKeyCredential
from azure.search.documents.indexes import SearchIndexClient
from azure.search.documents.indexes.models import (
    SearchIndex,
    SimpleField,
    SearchableField,
    VectorSearch,
    HnswAlgorithmConfiguration
)

# 初始化索引客户端
index_client = SearchIndexClient(
    endpoint="你的搜索服务端点",
    credential=AzureKeyCredential("你的API密钥")
)

# 配置向量搜索算法
vector_search = VectorSearch(
    algorithm_configurations=[
        HnswAlgorithmConfiguration(
            name="hnsw-config",
            kind="hnsw",
            parameters={"m": 4, "efConstruction": 400, "efSearch": 500}
        )
    ]
)

# 定义字段,重点注意embedding字段的类型
fields = [
    SimpleField(name="id", type="Edm.String", key=True, retrievable=True),
    SearchableField(name="content", type="Edm.String"),
    # 正确的向量字段定义
    SimpleField(
        name="embedding",
        type="Collection(Edm.Single)",
        searchable=True,
        retrievable=True,
        vector_search_configuration="hnsw-config"
    )
]

# 创建或更新索引
index = SearchIndex(name="你的索引名", fields=fields, vector_search=vector_search)
index_client.create_or_update_index(index)

2. 确保上传文档的向量值为数值数组

上传的embedding字段必须是浮点数组成的数组,不能是字符串或字符串数组。如果你的向量是从模型返回的字符串格式,需要先转换为数值数组:

from azure.search.documents import SearchClient

# 初始化搜索客户端
search_client = SearchClient(
    endpoint="你的搜索服务端点",
    index_name="你的索引名",
    credential=AzureKeyCredential("你的API密钥")
)

# 示例:正确的文档格式(embedding是数值数组)
documents = [
    {
        "id": "doc1",
        "content": "测试文档内容",
        "embedding": [0.0040753875, 0.12345, 0.6789]
    }
]

# 上传文档
result = search_client.upload_documents(documents=documents)
print(f"成功上传 {len(result)} 份文档")

如果你的向量原本是字符串(比如"0.0040753875,0.12345,0.6789"),可以用以下方式转换:

embedding_str = "0.0040753875,0.12345,0.6789"
embedding = list(map(float, embedding_str.split(',')))

关键检查点

  • 确认索引中embedding字段的type是Collection(Edm.Single)
  • 确认上传文档的embedding值是浮点数数组,而非字符串或字符串数组

内容的提问来源于stack exchange,提问作者Manu Chadha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 20:32:46