You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Azure AI Search中通过Split Skill获取分块索引

解决Azure AI Search分块记录ID索引问题

要给拆分后的每个分块添加记录ID字段并投影到索引中,需要修改技能集,添加ShaperSkill生成分块标识,再调整索引投影映射。具体操作如下:

步骤1:新增ShaperSkill生成分块ID

在现有技能集的skills数组中添加ShaperSkill,为每个pages元素附加chunk_id(即你需要的recordId),值为该分块在数组中的索引位置。修改后的技能数组:

"skills": [
  {
    "@odata.type": "#Microsoft.Skills.Text.SplitSkill",
    "name": "#3",
    "description": "Split skill to chunk documents",
    "context": "/document",
    "inputs": [
      {
        "name": "text",
        "source": "/document/content",
        "inputs": []
      }
    ],
    "outputs": [
      {
        "name": "textItems",
        "targetName": "pages"
      }
    ],
    "defaultLanguageCode": "en",
    "textSplitMode": "pages",
    "maximumPageLength": 2000,
    "pageOverlapLength": 500,
    "unit": "characters"
  },
  {
    "@odata.type": "#Microsoft.Skills.Util.ShaperSkill",
    "name": "#4",
    "context": "/document/pages/*",
    "inputs": [
      {
        "name": "text",
        "source": "/document/pages/*"
      },
      {
        "name": "chunk_id",
        "source": "$index"
      }
    ],
    "outputs": [
      {
        "name": "output",
        "targetName": "page_with_id"
      }
    ]
  }
]

说明:

  • $index是Azure AI Search内置变量,代表当前元素在父数组中的索引(从0开始),正好对应分块的位置标识。
  • ShaperSkill会将分块文本和chunk_id组合成新对象,存储在page_with_id字段中。

步骤2:调整索引投影映射

修改indexProjections的sourceContext和mappings,指向新生成的page_with_id字段,并添加recordId的映射规则:

"indexProjections": {
  "selectors": [
    {
      "targetIndexName": "testing-phase-1-docs-index",
      "parentKeyFieldName": "parent_id",
      "sourceContext": "/document/pages/*/page_with_id",
      "mappings": [
        {
          "name": "content",
          "source": "/document/pages/*/page_with_id/text"
        },
        {
          "name": "recordId",
          "source": "/document/pages/*/page_with_id/chunk_id"
        },
        {
          "name": "metadata_title",
          "source": "/document/metadata_title"
        }
      ]
    }
  ],
  "parameters": {
    "projectionMode": "skipIndexingParentDocuments"
  }
}

最终完整技能集配置

替换后的完整技能集JSON:

{
  "name": "testing-phase-1-docs-skillset",
  "description": "Skillset to chunk documents and generate embeddings",
  "skills": [
    {
      "@odata.type": "#Microsoft.Skills.Text.SplitSkill",
      "name": "#3",
      "description": "Split skill to chunk documents",
      "context": "/document",
      "inputs": [
        {
          "name": "text",
          "source": "/document/content",
          "inputs": []
        }
      ],
      "outputs": [
        {
          "name": "textItems",
          "targetName": "pages"
        }
      ],
      "defaultLanguageCode": "en",
      "textSplitMode": "pages",
      "maximumPageLength": 2000,
      "pageOverlapLength": 500,
      "unit": "characters"
    },
    {
      "@odata.type": "#Microsoft.Skills.Util.ShaperSkill",
      "name": "#4",
      "context": "/document/pages/*",
      "inputs": [
        {
          "name": "text",
          "source": "/document/pages/*"
        },
        {
          "name": "chunk_id",
          "source": "$index"
        }
      ],
      "outputs": [
        {
          "name": "output",
          "targetName": "page_with_id"
        }
      ]
    }
  ],
  "@odata.etag": "\"0x8DD029DA50735BD\"",
  "indexProjections": {
    "selectors": [
      {
        "targetIndexName": "testing-phase-1-docs-index",
        "parentKeyFieldName": "parent_id",
        "sourceContext": "/document/pages/*/page_with_id",
        "mappings": [
          {
            "name": "content",
            "source": "/document/pages/*/page_with_id/text"
          },
          {
            "name": "recordId",
            "source": "/document/pages/*/page_with_id/chunk_id"
          },
          {
            "name": "metadata_title",
            "source": "/document/metadata_title"
          }
        ]
      }
    ],
    "parameters": {
      "projectionMode": "skipIndexingParentDocuments"
    }
  }
}

注意事项

  • 确保目标索引testing-phase-1-docs-index已创建recordId字段(类型建议为Edm.Int32或Edm.String)。
  • 若需要自定义格式的ID,可在ShaperSkill中结合文档原始ID等字段拼接生成。

内容的提问来源于stack exchange,提问作者Yafaa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 09:05:09