自定义技能输出映射Edm.ComplexType失败,求正确方案
问题描述
我尝试了三种方法,试图将自定义技能的输出映射到搜索索引的Edm.ComplexType类型字段chunk_object中,但均未成功填充该字段。需求是搜索索引中的每个文档都包含chunk_object字段。
索引字段定义
{"name": "chunk_object", "type": "Edm.ComplexType", "fields": [ { "name": "chunk_content", "type": "Edm.String", "searchable": true, "filterable": true, "retrievable": true, "stored": true, "sortable": true, "facetable": true, "key": false, "indexAnalyzer": null, "searchAnalyzer": null, "analyzer": "standard.lucene", "normalizer": null, "dimensions": null, "vectorSearchProfile": null, "vectorEncoding": null, "synonymMaps": [] }, { "name": "page_start", "type": "Edm.Int64", "searchable": false, "filterable": true, "retrievable": true, "stored": true, "sortable": true, "facetable": true, "key": false, "indexAnalyzer": null, "searchAnalyzer": null, "analyzer": null, "normalizer": null, "dimensions": null, "vectorSearchProfile": null, "vectorEncoding": null, "synonymMaps": [] }, { "name": "page_end", "type": "Edm.Int64", "searchable": false, "filterable": true, "retrievable": true, "stored": true, "sortable": true, "facetable": true, "key": false, "indexAnalyzer": null, "searchAnalyzer": null, "analyzer": null, "normalizer": null, "dimensions": null, "vectorSearchProfile": null, "vectorEncoding": null, "synonymMaps": [] }, { "name": "chunk_idx", "type": "Edm.String", "searchable": true, "filterable": true, "retrievable": true, "stored": true, "sortable": true, "facetable": true, "key": false, "indexAnalyzer": null, "searchAnalyzer": null, "analyzer": "standard.lucene", "normalizer": null, "dimensions": null, "vectorSearchProfile": null, "vectorEncoding": null, "synonymMaps": [] } ] }
自定义技能输出
自定义技能的输出映射到/document/jsonChunks/*,包含239个对象,输出结构如下:
{ "values": [ { "recordId": "1", "data": { "jsonChunks": [ { "chunk": "this is chunk 1", "page_start": 1, "page_end": 1, "chunk_idx": "#1-file.pdf'" }, { "chunk": "this is chunk 2", "page_start": 1, "page_end": 1, "chunk_idx": "#1-file.pdf'" } ] } } ] }
内存输出结构:
-/document/jsonChunks Object[239] -/* -/chunk_content -/page_start -/page_end -/chunk_idx
尝试的三种Shaper Skill配置方法
方法1
{ "@odata.type": "#Microsoft.Skills.Util.ShaperSkill", "name": "#2", "description": "", "context": "/document", "inputs": [ { "name": "chunk_content", "source": "/document/jsonChunks/*/chunk" }, { "name": "page_start", "source": "/document/jsonChunks/*/page_start" }, { "name": "page_end", "source": "/document/jsonChunks/*/page_end" }, { "name": "chunk_idx", "source": "/document/jsonChunks/*/chunk_idx" } ], "outputs": [ { "name": "output", "targetName": "chunk_object" } ] }
对应的内存输出:
/document/chunk_object Object -/chunk_content Object[239] -/* -/page_start Object[239] -/* -/page_end Object[239] -/* -/chunk_idx Object[239] -/*
方法2
{ "@odata.type": "#Microsoft.Skills.Util.ShaperSkill", "name": "#2", "description": "", "context": "/document", "inputs": [ { "name": "jsonChunk", "source": "/document/jsonChunks/*" } ], "outputs": [ { "name": "output", "targetName": "chunk_object" } ] }
对应的内存输出:
/document/chunk_object Object -/jsonChunk Object[239] -/* -/chunk_content -/page_start -/page_end -/chunk_idx
方法3
{ "@odata.type": "#Microsoft.Skills.Util.ShaperSkill", "name": "#2", "description": "", "context": "/document", "inputs": [ { "name": "jsonChunk", "sourceContext": "/document/jsonChunks/*", "inputs": [ { "name": "chunk_object", "source": "/document/jsonChunks/*/chunk" }, { "name": "page_start", "source": "/document/jsonChunks/*/page_start" }, { "name": "page_end", "source": "/document/jsonChunks/*/page_end" }, { "name": "chunk_idx", "source": "/document/jsonChunks/*/chunk_idx" } ] } ], "outputs": [ { "name": "output", "targetName": "chunk_object" } ] }
对应的内存输出:
/document/chunk_object Object -/jsonChunk Object[239] -/* -/chunk_content -/page_start -/page_end -/chunk_idx
以上三种方法均未成功填充索引字段,需要正确的实现方法。
解决方案
1. 修正索引字段类型
自定义技能输出是多个chunk对象,当前索引中chunk_object的类型是Edm.ComplexType(单个复杂对象),无法存储集合数据。需要将其修改为Collection(Edm.ComplexType):
{"name": "chunk_object", "type": "Collection(Edm.ComplexType)", "fields": [ // 子字段定义保持不变 { "name": "chunk_content", "type": "Edm.String", "searchable": true, "filterable": true, "retrievable": true, "stored": true, "sortable": true, "facetable": true, "key": false, "indexAnalyzer": null, "searchAnalyzer": null, "analyzer": "standard.lucene", "normalizer": null, "dimensions": null, "vectorSearchProfile": null, "vectorEncoding": null, "synonymMaps": [] }, { "name": "page_start", "type": "Edm.Int64", "searchable": false, "filterable": true, "retrievable": true, "stored": true, "sortable": true, "facetable": true, "key": false, "indexAnalyzer": null, "searchAnalyzer": null, "analyzer": null, "normalizer": null, "dimensions": null, "vectorSearchProfile": null, "vectorEncoding": null, "synonymMaps": [] }, { "name": "page_end", "type": "Edm.Int64", "searchable": false, "filterable": true, "retrievable": true, "stored": true, "sortable": true, "facetable": true, "key": false, "indexAnalyzer": null, "searchAnalyzer": null, "analyzer": null, "normalizer": null, "dimensions": null, "vectorSearchProfile": null, "vectorEncoding": null, "synonymMaps": [] }, { "name": "chunk_idx", "type": "Edm.String", "searchable": true, "filterable": true, "retrievable": true, "stored": true, "sortable": true, "facetable": true, "key": false, "indexAnalyzer": null, "searchAnalyzer": null, "analyzer": "standard.lucene", "normalizer": null, "dimensions": null, "vectorSearchProfile": null, "vectorEncoding": null, "synonymMaps": [] } ] }
2. 正确配置Shaper Skill
使用嵌套sourceContext的方式,遍历每个chunk并构建符合chunk_object结构的对象,最终生成集合数组:
{ "@odata.type": "#Microsoft.Skills.Util.ShaperSkill", "name": "shape-chunk-collection", "description": "Convert each jsonChunk to chunk_object and collect into a collection", "context": "/document", "inputs": [ { "name": "chunk_object", "sourceContext": "/document/jsonChunks/*", "inputs": [ { "name": "chunk_content", "source": "/chunk" }, { "name": "page_start", "source": "/page_start" }, { "name": "page_end", "source": "/page_end" }, { "name": "chunk_idx", "source": "/chunk_idx" } ] } ], "outputs": [ { "name": "output", "targetName": "chunk_object" } ] }
配置说明
sourceContext: "/document/jsonChunks/*":指定遍历jsonChunks数组中的每个元素- 内部
inputs:为每个元素提取字段,source使用相对路径(基于当前遍历的元素,无需完整路径) - 最终输出的
/document/chunk_object是包含所有符合结构的chunk对象的数组,匹配索引中Collection(Edm.ComplexType)类型的字段
3. 索引字段映射
确保在索引的字段映射中,将/document/chunk_object直接映射到索引的chunk_object字段即可。
内容的提问来源于stack exchange,提问作者MajorMajorMajorMajor
相关产品推荐
相关产品推荐

