You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义技能输出映射Edm.ComplexType失败,求正确方案

问题描述

我尝试了三种方法,试图将自定义技能的输出映射到搜索索引的Edm.ComplexType类型字段chunk_object中,但均未成功填充该字段。需求是搜索索引中的每个文档都包含chunk_object字段。

索引字段定义

{"name": "chunk_object",
      "type": "Edm.ComplexType",
      "fields": [
        {
          "name": "chunk_content",
          "type": "Edm.String",
          "searchable": true,
          "filterable": true,
          "retrievable": true,
          "stored": true,
          "sortable": true,
          "facetable": true,
          "key": false,
          "indexAnalyzer": null,
          "searchAnalyzer": null,
          "analyzer": "standard.lucene",
          "normalizer": null,
          "dimensions": null,
          "vectorSearchProfile": null,
          "vectorEncoding": null,
          "synonymMaps": []
        },
        {
          "name": "page_start",
          "type": "Edm.Int64",
          "searchable": false,
          "filterable": true,
          "retrievable": true,
          "stored": true,
          "sortable": true,
          "facetable": true,
          "key": false,
          "indexAnalyzer": null,
          "searchAnalyzer": null,
          "analyzer": null,
          "normalizer": null,
          "dimensions": null,
          "vectorSearchProfile": null,
          "vectorEncoding": null,
          "synonymMaps": []
        },
        {
          "name": "page_end",
          "type": "Edm.Int64",
          "searchable": false,
          "filterable": true,
          "retrievable": true,
          "stored": true,
          "sortable": true,
          "facetable": true,
          "key": false,
          "indexAnalyzer": null,
          "searchAnalyzer": null,
          "analyzer": null,
          "normalizer": null,
          "dimensions": null,
          "vectorSearchProfile": null,
          "vectorEncoding": null,
          "synonymMaps": []
        },
        {
          "name": "chunk_idx",
          "type": "Edm.String",
          "searchable": true,
          "filterable": true,
          "retrievable": true,
          "stored": true,
          "sortable": true,
          "facetable": true,
          "key": false,
          "indexAnalyzer": null,
          "searchAnalyzer": null,
          "analyzer": "standard.lucene",
          "normalizer": null,
          "dimensions": null,
          "vectorSearchProfile": null,
          "vectorEncoding": null,
          "synonymMaps": []
        }
      ]
}

自定义技能输出

自定义技能的输出映射到/document/jsonChunks/*,包含239个对象,输出结构如下:

{
    "values": [
        {
            "recordId": "1",
            "data": {
                "jsonChunks": [
                    {
                        "chunk": "this is chunk 1",
                        "page_start": 1,
                        "page_end": 1,
                        "chunk_idx": "#1-file.pdf'"
                    },
                    {
                        "chunk": "this is chunk 2",
                        "page_start": 1,
                        "page_end": 1,
                        "chunk_idx": "#1-file.pdf'"
                    }
                ]
            }
        }
    ]
}

内存输出结构:

-/document/jsonChunks Object[239]
  -/*
    -/chunk_content
    -/page_start
    -/page_end
    -/chunk_idx

尝试的三种Shaper Skill配置方法

方法1

{
    "@odata.type": "#Microsoft.Skills.Util.ShaperSkill",
    "name": "#2",
    "description": "",
    "context": "/document",
    "inputs": [
        {
            "name": "chunk_content",
            "source": "/document/jsonChunks/*/chunk"
        },
        {
            "name": "page_start",
            "source": "/document/jsonChunks/*/page_start"
        },
        {
            "name": "page_end",
            "source": "/document/jsonChunks/*/page_end"
        },
        {
            "name": "chunk_idx",
            "source": "/document/jsonChunks/*/chunk_idx"
        }
    ],
    "outputs": [
        {
            "name": "output",
            "targetName": "chunk_object"
        }
    ]
}

对应的内存输出:

/document/chunk_object Object
  -/chunk_content Object[239]
    -/*
  -/page_start Object[239]
    -/*
  -/page_end Object[239]
    -/*
  -/chunk_idx Object[239]
    -/*

方法2

{
  "@odata.type": "#Microsoft.Skills.Util.ShaperSkill",
  "name": "#2",
  "description": "",
  "context": "/document",
  "inputs": [
    {
      "name": "jsonChunk",
      "source": "/document/jsonChunks/*"
    }
  ],
  "outputs": [
    {
      "name": "output",
      "targetName": "chunk_object"
    }
  ]
}

对应的内存输出:

/document/chunk_object Object
  -/jsonChunk Object[239]
    -/*
      -/chunk_content
      -/page_start
      -/page_end
      -/chunk_idx

方法3

{
  "@odata.type": "#Microsoft.Skills.Util.ShaperSkill",
  "name": "#2",
  "description": "",
  "context": "/document",
  "inputs": [
    {
      "name": "jsonChunk",
      "sourceContext": "/document/jsonChunks/*",
      "inputs": [
        {
          "name": "chunk_object",
          "source": "/document/jsonChunks/*/chunk"
        },
        {
          "name": "page_start",
          "source": "/document/jsonChunks/*/page_start"
        },
        {
          "name": "page_end",
          "source": "/document/jsonChunks/*/page_end"
        },
        {
          "name": "chunk_idx",
          "source": "/document/jsonChunks/*/chunk_idx"
        }
      ]
    }
  ],
  "outputs": [
    {
      "name": "output",
      "targetName": "chunk_object"
    }
  ]
}

对应的内存输出:

/document/chunk_object Object
  -/jsonChunk Object[239]
    -/*
      -/chunk_content
      -/page_start
      -/page_end
      -/chunk_idx

以上三种方法均未成功填充索引字段,需要正确的实现方法。

解决方案

1. 修正索引字段类型

自定义技能输出是多个chunk对象,当前索引中chunk_object的类型是Edm.ComplexType(单个复杂对象),无法存储集合数据。需要将其修改为Collection(Edm.ComplexType):

{"name": "chunk_object",
      "type": "Collection(Edm.ComplexType)",
      "fields": [
        // 子字段定义保持不变
        {
          "name": "chunk_content",
          "type": "Edm.String",
          "searchable": true,
          "filterable": true,
          "retrievable": true,
          "stored": true,
          "sortable": true,
          "facetable": true,
          "key": false,
          "indexAnalyzer": null,
          "searchAnalyzer": null,
          "analyzer": "standard.lucene",
          "normalizer": null,
          "dimensions": null,
          "vectorSearchProfile": null,
          "vectorEncoding": null,
          "synonymMaps": []
        },
        {
          "name": "page_start",
          "type": "Edm.Int64",
          "searchable": false,
          "filterable": true,
          "retrievable": true,
          "stored": true,
          "sortable": true,
          "facetable": true,
          "key": false,
          "indexAnalyzer": null,
          "searchAnalyzer": null,
          "analyzer": null,
          "normalizer": null,
          "dimensions": null,
          "vectorSearchProfile": null,
          "vectorEncoding": null,
          "synonymMaps": []
        },
        {
          "name": "page_end",
          "type": "Edm.Int64",
          "searchable": false,
          "filterable": true,
          "retrievable": true,
          "stored": true,
          "sortable": true,
          "facetable": true,
          "key": false,
          "indexAnalyzer": null,
          "searchAnalyzer": null,
          "analyzer": null,
          "normalizer": null,
          "dimensions": null,
          "vectorSearchProfile": null,
          "vectorEncoding": null,
          "synonymMaps": []
        },
        {
          "name": "chunk_idx",
          "type": "Edm.String",
          "searchable": true,
          "filterable": true,
          "retrievable": true,
          "stored": true,
          "sortable": true,
          "facetable": true,
          "key": false,
          "indexAnalyzer": null,
          "searchAnalyzer": null,
          "analyzer": "standard.lucene",
          "normalizer": null,
          "dimensions": null,
          "vectorSearchProfile": null,
          "vectorEncoding": null,
          "synonymMaps": []
        }
      ]
}

2. 正确配置Shaper Skill

使用嵌套sourceContext的方式,遍历每个chunk并构建符合chunk_object结构的对象,最终生成集合数组:

{
  "@odata.type": "#Microsoft.Skills.Util.ShaperSkill",
  "name": "shape-chunk-collection",
  "description": "Convert each jsonChunk to chunk_object and collect into a collection",
  "context": "/document",
  "inputs": [
    {
      "name": "chunk_object",
      "sourceContext": "/document/jsonChunks/*",
      "inputs": [
        {
          "name": "chunk_content",
          "source": "/chunk"
        },
        {
          "name": "page_start",
          "source": "/page_start"
        },
        {
          "name": "page_end",
          "source": "/page_end"
        },
        {
          "name": "chunk_idx",
          "source": "/chunk_idx"
        }
      ]
    }
  ],
  "outputs": [
    {
      "name": "output",
      "targetName": "chunk_object"
    }
  ]
}

配置说明

  • sourceContext: "/document/jsonChunks/*":指定遍历jsonChunks数组中的每个元素
  • 内部inputs:为每个元素提取字段,source使用相对路径(基于当前遍历的元素,无需完整路径)
  • 最终输出的/document/chunk_object是包含所有符合结构的chunk对象的数组,匹配索引中Collection(Edm.ComplexType)类型的字段

3. 索引字段映射

确保在索引的字段映射中,将/document/chunk_object直接映射到索引的chunk_object字段即可。

内容的提问来源于stack exchange,提问作者MajorMajorMajorMajor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 03:29:56