Azure添加向量嵌入到索引时出现技能集错误求助
Azure搜索技能集向量嵌入错误修复
问题根源
报错提示Cannot iterate over non-array '/document/pages'和输入类型不匹配,核心原因是AzureOpenAIEmbeddingSkill的上下文配置错误:
- SplitSkill已将文档拆分为数组类型的
/document/pages,但EmbeddingSkill的context设为/document(文档根级别),此时使用/document/pages/*会尝试在根级别迭代数组,不符合技能对单个字符串输入的要求。
修复方案
调整EmbeddingSkill的上下文到数组项级别,让技能逐个处理拆分后的每个page:
- 将EmbeddingSkill的
context改为/document/pages/*,表示遍历/document/pages数组中的每一项 - 输入source简化为
/(当前上下文即为单个page的内容),或保持/document/pages/*(配合新上下文也可正常工作)
修改后的技能集配置
{ "@odata.context": "https://redacted/$metadata#skillsets/$entity", "@odata.etag": "\"something\"", "name": "something-skillset", "description": "", "skills": [ { "@odata.type": "#Microsoft.Skills.Text.SplitSkill", "name": "Text split skill", "description": "Splits text into pages small enough to vectorize", "context": "/document", "defaultLanguageCode": "en", "textSplitMode": "pages", "maximumPageLength": 2000, "pageOverlapLength": 500, "maximumPagesToTake": 0, "inputs": [ { "name": "text", "source": "/document/content" } ], "outputs": [ { "name": "textItems", "targetName": "/document/pages" } ] }, { "@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill", "name": "Create vector embedding for pages", "description": "", "context": "/document/pages/*", "resourceUri": "https://something.openai.azure.com", "apiKey": "<redacted>", "deploymentId": "text-embedding-ada-002", "inputs": [ { "name": "text", "source": "/" } ], "outputs": [ { "name": "embedding", "targetName": "contentVector" } ], "authIdentity": null } ], "cognitiveServices": { "@odata.type": "#Microsoft.Azure.Search.DefaultCognitiveServices", "description": null }, "knowledgeStore": null, "indexProjections": null, "encryptionKey": null }
额外说明
- 调整后,EmbeddingSkill会自动遍历
/document/pages中的每个字符串元素,为每个page生成对应的向量嵌入 - 若要将向量映射到索引字段,需确保索引中配置了类型为
Collection(Edm.Single)的向量字段,并在索引器的字段映射中关联/document/pages/*/contentVector
内容的提问来源于stack exchange,提问作者Toodleey
相关产品推荐
相关产品推荐

