创建Azure AI Search索引时遇向量维度不匹配问题求助
Azure AI Search向量字段长度为0错误排查与解决
问题
创建Azure AI Search索引时,Blob存储中的JSON文档在调试时报错:向量字段'vector'维度为1536,期望长度1536,但提供的向量长度为0。已验证文本拆分器输入/document/combined_text非空,但怀疑拆分器输出存在空文本块。
技能集原始配置
"skills": [ { "@odata.type": "#Microsoft.Skills.Text.SplitSkill", "name": "SplitTextSkill", "description": "Split combined text into chunks of 4096 characters", "context": "/document", "defaultLanguageCode": "en", "textSplitMode": "pages", "maximumPageLength": 4096, "pageOverlapLength": 0, "maximumPagesToTake": 0, "inputs": [ { "name": "text", "source": "/document/combined_text" } ], "outputs": [ { "name": "textItems", "targetName": "chunks" } ] }, { "@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill", "name": "Embeddings generation", "description": "Azure OpenAI Embedding Skill", "context": "/document/chunks/*", "resourceUri": "endpointOfAzureOpenaiService", "apiKey": "<redacted>", "deploymentId": "akm-aml-embeddings", "dimensions": 1536, "modelName": "text-embedding-ada-002", "inputs": [ { "name": "text", "source": "/document/chunks/*" } ], "outputs": [ { "name": "embedding", "targetName": "vector" } ], "authIdentity": null } ]
解决方案
1. 过滤空文本块
拆分技能可能生成空字符串块(比如原文本含大量空白字符、换行符),导致嵌入技能输出长度为0的向量。在拆分技能和嵌入技能之间添加条件筛选技能,只保留非空chunk:
{ "@odata.type": "#Microsoft.Skills.Util.ConditionalSkill", "name": "FilterEmptyChunks", "context": "/document", "inputs": [ { "name": "condition", "source": "@not(empty(/document/chunks))" }, { "name": "whenTrue", "source": "/document/chunks" }, { "name": "whenFalse", "source": "[]" } ], "outputs": [ { "name": "output", "targetName": "filteredChunks" } ] }
随后修改嵌入技能的上下文和输入源:
{ "@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill", "name": "Embeddings generation", "context": "/document/filteredChunks/*", // 其他配置保持不变 "inputs": [ { "name": "text", "source": "/document/filteredChunks/*" } ] }
2. 优化拆分器配置
- 将
textSplitMode从pages改为sentences,避免分页逻辑产生空块; - 若使用的API版本支持,添加
minimumPageLength参数,确保生成的chunk长度不低于阈值; - 预处理输入文本,去除多余空白字符:添加文本合并技能清理
combined_text:
{ "@odata.type": "#Microsoft.Skills.Text.MergeSkill", "name": "CleanCombinedText", "context": "/document", "inputs": [ { "name": "text", "source": "/document/combined_text" }, { "name": "insertPreTag", "source": "' '" }, { "name": "insertPostTag", "source": "' '" } ], "outputs": [ { "name": "mergedText", "targetName": "cleaned_text" } ] }
之后将拆分技能的输入源改为/document/cleaned_text。
3. 验证嵌入技能输入
确认嵌入技能的输入/document/chunks/*指向单个文本块内容(拆分技能输出的textItems是字符串数组,每个元素为文本块,此配置正确),但需确保无空字符串元素。
总结
空文本块是导致向量长度为0的核心原因,通过过滤空块、优化拆分配置或预处理文本,即可解决索引维度校验错误。
内容的提问来源于stack exchange,提问作者user25708898
相关产品推荐
相关产品推荐

