如何在Azure AI Search中通过Split Skill获取分块索引
解决Azure AI Search分块记录ID索引问题
要给拆分后的每个分块添加记录ID字段并投影到索引中,需要修改技能集,添加ShaperSkill生成分块标识,再调整索引投影映射。具体操作如下:
步骤1:新增ShaperSkill生成分块ID
在现有技能集的skills数组中添加ShaperSkill,为每个pages元素附加chunk_id(即你需要的recordId),值为该分块在数组中的索引位置。修改后的技能数组:
"skills": [ { "@odata.type": "#Microsoft.Skills.Text.SplitSkill", "name": "#3", "description": "Split skill to chunk documents", "context": "/document", "inputs": [ { "name": "text", "source": "/document/content", "inputs": [] } ], "outputs": [ { "name": "textItems", "targetName": "pages" } ], "defaultLanguageCode": "en", "textSplitMode": "pages", "maximumPageLength": 2000, "pageOverlapLength": 500, "unit": "characters" }, { "@odata.type": "#Microsoft.Skills.Util.ShaperSkill", "name": "#4", "context": "/document/pages/*", "inputs": [ { "name": "text", "source": "/document/pages/*" }, { "name": "chunk_id", "source": "$index" } ], "outputs": [ { "name": "output", "targetName": "page_with_id" } ] } ]
说明:
$index是Azure AI Search内置变量,代表当前元素在父数组中的索引(从0开始),正好对应分块的位置标识。- ShaperSkill会将分块文本和
chunk_id组合成新对象,存储在page_with_id字段中。
步骤2:调整索引投影映射
修改indexProjections的sourceContext和mappings,指向新生成的page_with_id字段,并添加recordId的映射规则:
"indexProjections": { "selectors": [ { "targetIndexName": "testing-phase-1-docs-index", "parentKeyFieldName": "parent_id", "sourceContext": "/document/pages/*/page_with_id", "mappings": [ { "name": "content", "source": "/document/pages/*/page_with_id/text" }, { "name": "recordId", "source": "/document/pages/*/page_with_id/chunk_id" }, { "name": "metadata_title", "source": "/document/metadata_title" } ] } ], "parameters": { "projectionMode": "skipIndexingParentDocuments" } }
最终完整技能集配置
替换后的完整技能集JSON:
{ "name": "testing-phase-1-docs-skillset", "description": "Skillset to chunk documents and generate embeddings", "skills": [ { "@odata.type": "#Microsoft.Skills.Text.SplitSkill", "name": "#3", "description": "Split skill to chunk documents", "context": "/document", "inputs": [ { "name": "text", "source": "/document/content", "inputs": [] } ], "outputs": [ { "name": "textItems", "targetName": "pages" } ], "defaultLanguageCode": "en", "textSplitMode": "pages", "maximumPageLength": 2000, "pageOverlapLength": 500, "unit": "characters" }, { "@odata.type": "#Microsoft.Skills.Util.ShaperSkill", "name": "#4", "context": "/document/pages/*", "inputs": [ { "name": "text", "source": "/document/pages/*" }, { "name": "chunk_id", "source": "$index" } ], "outputs": [ { "name": "output", "targetName": "page_with_id" } ] } ], "@odata.etag": "\"0x8DD029DA50735BD\"", "indexProjections": { "selectors": [ { "targetIndexName": "testing-phase-1-docs-index", "parentKeyFieldName": "parent_id", "sourceContext": "/document/pages/*/page_with_id", "mappings": [ { "name": "content", "source": "/document/pages/*/page_with_id/text" }, { "name": "recordId", "source": "/document/pages/*/page_with_id/chunk_id" }, { "name": "metadata_title", "source": "/document/metadata_title" } ] } ], "parameters": { "projectionMode": "skipIndexingParentDocuments" } } }
注意事项
- 确保目标索引
testing-phase-1-docs-index已创建recordId字段(类型建议为Edm.Int32或Edm.String)。 - 若需要自定义格式的ID,可在ShaperSkill中结合文档原始ID等字段拼接生成。
内容的提问来源于stack exchange,提问作者Yafaa
相关产品推荐
相关产品推荐

