如何在Azure AI Search索引中映射TextSplitSkill的ordinalPositions?
解决Azure AI Search中TextSplitSkill的ordinalPositions投影问题
可以用TextSplitSkill + ShaperSkill + indexProjections实现需求
核心是通过ShaperSkill将两个数组的对应元素配对成单个对象,再通过索引投影的$each语法逐个映射到索引字段,避免直接映射整个数组导致的序列化问题。
步骤1:正确配置TextSplitSkill
确保技能输出拆分后的文本块数组textItems和对应顺序的序号数组ordinalPositions:
{ "@odata.type": "#Microsoft.Skills.Text.SplitSkill", "name": "text-split", "description": "拆分文本为块", "context": "/document", "defaultLanguageCode": "zh-CN", "textSplitMode": "pages", "maximumPageLength": 1000, "inputs": [ { "name": "text", "source": "/document/content" } ], "outputs": [ { "name": "textItems", "targetName": "chunks" }, { "name": "ordinalPositions", "targetName": "chunk_orders" } ] }
步骤2:用ShaperSkill配对数组元素
利用ShaperSkill的$index变量,将chunks和chunk_orders数组中同索引的元素组合成单个对象,生成包含文本块与对应序号的对象数组:
{ "@odata.type": "#Microsoft.Skills.Util.ShaperSkill", "name": "shape-chunk-with-order", "context": "/document", "inputs": [ { "name": "content", "source": "/document/chunks/*" }, { "name": "chunk_index", "source": "/document/chunk_orders/*", "sourceContext": "/document/chunks/$index" } ], "outputs": [ { "name": "output", "targetName": "chunk_with_order" } ] }
步骤3:配置索引投影
在索引投影中,使用$each遍历chunk_with_order数组,将每个对象的字段映射到索引对应字段(注意索引的chunk_index需设为整数类型):
"indexProjections": { "selectors": [ { "targetIndexName": "你的目标索引名", "parentKeyFieldName": "document_id", "sourceContext": "/document/chunk_with_order/*", "mappings": [ { "source": "/document/chunk_with_order/*/content", "targetName": "chunk_content" }, { "source": "/document/chunk_with_order/*/chunk_index", "targetName": "chunk_index" } ] } ], "parameters": { "projectionMode": "skipIndexingParentDocument" } }
若上述方案不可行,ordinalPositions的其他使用场景
- 查询后重组文档:从索引检索到多个文本块时,通过
ordinalPositions的值对chunk排序,还原原始文档的顺序。 - 技能链内排序:在后续技能处理中,利用序号对文本块排序,确保处理逻辑符合原始文档的顺序。
- 元数据关联:将序号作为chunk的元数据,用于统计文档拆分后的块数量、定位特定位置的chunk等。
内容的提问来源于stack exchange,提问作者Daniel Georgiev
相关产品推荐
相关产品推荐

