You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure AI Search问题:索引分块文档时向量维度不匹配

Azure AI Search RAG系统分块文档索引维度不匹配问题

我的设置概述

  • 通过SharePoint导入数据(此环节正常)
  • 将文件分块为页面
  • 使用ada-002模型为每个页面生成嵌入向量
  • 生成的嵌入向量传入索引用于搜索

问题

尝试将分块页面传入索引器时,偶尔出现维度不匹配错误:

There's a mismatch in vector dimensions. The vector field 'content_embeddings', with dimension of '1536', expects a length of '1536'. However, the provided vector has a length of '3072'. Please ensure that the vector length matches the expected length of the vector field.

观察结果

  • 文档未分块(大小小于分块限制)时,可成功索引
  • 问题仅在文档分块时出现

已采取的排查步骤

  • 输出验证:记录嵌入后、索引前的向量维度,多数为正确的1536,但分块向量偶尔出现维度过大
  • 分块逻辑:确认分块过程未合并多个分块或重叠内容,问题仍存在
  • 索引器配置:检查索引器设置,确认其映射到content_embeddings字段

我的疑问

  • 分块嵌入向量索引时,维度不匹配的原因是什么?
  • Azure AI Search中,嵌入向量有哪些特定配置或最佳实践?

相关配置代码

技能集配置

"skills": [
    {
      "@odata.type": "#Microsoft.Skills.Text.SplitSkill",
      "name": "SplitSkill",
      "description": "A skill that splits text into chunks",
      "context": "/document",
      "defaultLanguageCode": "en",
      "textSplitMode": "pages",
      "maximumPageLength": 2000,
      "pageOverlapLength": 500,
      "maximumPagesToTake": 0,
      "unit": "azureOpenAITokens",
      "inputs": [
        {
          "name": "text",
          "source": "/document/content"
        }
      ],
      "outputs": [
        {
          "name": "textItems",
          "targetName": "pages"
        }
      ],
      "azureOpenAITokenizerParameters": {
        "encoderModelName": "cl100k_base",
        "allowedSpecialTokens": [
          "[START]",
          "[END]"
        ]
      }
    },
    {
      "@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill",
      "name": "ContentEmbeddingSkill",
      "description": "Connects to Azure OpenAI deployed embedding model to generate embeddings from content.",
      "context": "/document/pages/*",
      "resourceUri": "https://xxxx.openai.azure.com",
      "apiKey": "<redacted>",
      "deploymentId": "text-embedding-ada-002",
      "dimensions": 1536,
      "modelName": "text-embedding-ada-002",
      "inputs": [
        {
          "name": "text",
          "source": "/document/pages/*"
        }
      ],
      "outputs": [
        {
          "name": "embedding",
          "targetName": "content_embeddings"
        }
      ],
      "authIdentity": null
    }

索引器配置

{
  "@odata.context": "xxxxxxxxxxx",
  "@odata.etag": "xxxxxxxxxxx",
  "name": "xxxxxxxxxxx-vector",
  "description": null,
  "dataSourceName": "sharepoint-datasource",
  "skillsetName": "contentembedding",
  "targetIndexName": "sharepoint-index",
  "disabled": null,
  "schedule": null,
  "parameters": {
    "batchSize": 10,
    "maxFailedItems": 100,
    "maxFailedItemsPerBatch": null,
    "base64EncodeKeys": null,
    "configuration": {
      "indexedFileNameExtensions": ".csv, .docx, .pptx,.txt,.html,.pdf",
      "excludedFileNameExtensions": ".png, .jpg, .gif",
      "dataToExtract": "contentAndMetadata"
    }
  },
  "fieldMappings": [
    {
      "sourceFieldName": "content",
      "targetFieldName": "content",
      "mappingFunction": null
    }
  ],
  "outputFieldMappings": [
    {
      "sourceFieldName": "/document/pages",
      "targetFieldName": "pages",
      "mappingFunction": null
    },
    {
      "sourceFieldName": "/document/pages/*/content_embeddings/*",
      "targetFieldName": "content_embeddings",
      "mappingFunction": null
    }
  ],
  "cache": null,
  "encryptionKey": null
}

问题原因与解决建议

原因分析

  1. 索引器输出映射错误:当前outputFieldMappings中content_embeddings的源路径/document/pages/*/content_embeddings/*会将所有分块的嵌入向量扁平化拼接,当文档有2个分块时,就会生成1536×2=3072维度的向量,与索引字段的1536维度要求冲突。
  2. 未启用分块投影:默认情况下,索引器会将整个原始文档作为一个索引条目,分块后的内容和向量会被合并到该条目中,而非每个分块单独成为索引条目。

解决步骤

  1. 修正输出映射路径
    将content_embeddings的源路径改为/document/pages/*/content_embeddings,避免扁平化拼接向量:

    {
      "sourceFieldName": "/document/pages/*/content_embeddings",
      "targetFieldName": "content_embeddings",
      "mappingFunction": null
    }
    
  2. 配置分块投影(核心解决方法)
    在索引器配置中添加projections节点,让每个分块成为独立的索引条目,确保每个索引条目的content_embeddings为单个1536维度向量:

    "projections": [
      {
        "selectors": [
          {
            "source": "/document/pages/*",
            "targetName": "content"
          },
          {
            "source": "/document/pages/*/content_embeddings",
            "targetName": "content_embeddings"
          },
          // 保留原始文档元数据(可选)
          {
            "source": "/document/metadata_title",
            "targetName": "metadata_title"
          },
          {
            "source": "/document/metadata_path",
            "targetName": "metadata_path"
          }
        ],
        "tables": [
          {
            "tableName": "sharepoint-index",
            "keyFieldName": "id"
          }
        ]
      }
    ]
    

    注意:索引的id字段需支持自动生成,或通过映射生成包含分块标识的唯一ID(比如原始文档ID+分块序号)。

  3. 验证嵌入技能上下文
    当前嵌入技能的上下文/document/pages/*配置正确,确保每个分块单独生成嵌入向量,无需修改。

Azure AI Search嵌入向量最佳实践

  • 分块投影必选:处理长文档时必须启用分块投影,让每个分块成为独立索引条目,避免向量合并问题。
  • 字段类型匹配:索引中的向量字段需设置为Edm.Single数组(单向量,维度1536),若需存储多向量则用Collection(Edm.Single),但RAG场景通常每个分块对应一个单向量条目。
  • 模型维度一致:确保嵌入技能的dimensions参数与索引字段维度完全匹配,ada-002默认维度为1536,无需修改。
  • 日志监控:开启索引器日志,跟踪每个分块的向量生成情况,快速定位异常来源。

内容的提问来源于stack exchange,提问作者el Josso

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 02:22:03