Azure搜索索引器处理删除Blob时的警告消除方案咨询
解决Azure搜索索引器软删除时的投影警告问题
问题说明
通过Azure门户「导入并向量化数据」创建包含搜索实例、技能集和索引器的RAG环境,添加文档功能正常。启用软删除后删除容器中的文档,索引器可正确移除Blob及搜索服务中对应的拆分/嵌入文档,但控制台出现以下警告:
operation Projection.IndexProjections.SearchIndex.<indexname> Message Could not generate projection from input '/document/pages/*'. Check the 'source' or 'sourceContext' property of your projection in your skillset. =$(/document/pages/*) ?map { "chunk": $(/document/pages/*), "doc_name": $(/document/metadata_storage_name), "source": $(/document/owner), "title": $(/document/title), "vector": $(/document/pages/*/vector) }
当前使用的技能集配置如下:
{ "@odata.context": "https://<myservice>.search.windows.net/$metadata#skillsets/$entity", "@odata.etag": "\"**********\"", "name": "<my>-skillset", "description": "Skillset to chunk documents and generate embeddings", "skills": [ { "@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill", "name": "#1", "description": null, "context": "/document/pages/*", "resourceUri": "https://<myresource>.openai.azure.com", "apiKey": "<redacted>", "deploymentId": "text-embedding-ada-002", "inputs": [ { "name": "text", "source": "/document/pages/*" } ], "outputs": [ { "name": "embedding", "targetName": "vector" } ], "authIdentity": null }, { "@odata.type": "#Microsoft.Skills.Text.SplitSkill", "name": "#2", "description": "Split skill to chunk documents", "context": "/document", "defaultLanguageCode": "en", "textSplitMode": "pages", "maximumPageLength": 5000, "pageOverlapLength": 1250, "maximumPagesToTake": 0, "inputs": [ { "name": "text", "source": "/document/content" } ], "outputs": [ { "name": "textItems", "targetName": "pages" } ] } ], "cognitiveServices": null, "knowledgeStore": null, "indexProjections": { "selectors": [ { "targetIndexName": "<myindex>", "parentKeyFieldName": "parent_id", "sourceContext": "/document/pages/*", "mappings": [ { "name": "chunk", "source": "/document/pages/*", "sourceContext": null, "inputs": [] }, { "name": "vector", "source": "/document/pages/*/vector", "sourceContext": null, "inputs": [] }, { "name": "doc_name", "source": "/document/metadata_storage_name", "sourceContext": null, "inputs": [] }, { "name": "title", "source": "/document/title", "sourceContext": null, "inputs": [] }, { "name": "source", "source": "/document/owner", "sourceContext": null, "inputs": [] } ] } ], "parameters": { "projectionMode": "skipIndexingParentDocuments" } }, "encryptionKey": null }
抑制/规避警告的可行思路
- 调整索引器删除处理逻辑:在索引器的
deleteDetectionPolicy中,配置仅当检测到软删除标记时直接触发索引删除操作,跳过技能集的执行流程。可通过自定义字段映射或条件判断,避免删除操作进入投影环节。 - 为投影添加存在性条件判断:修改技能集
indexProjections中selectors的sourceContext,使用条件表达式过滤掉无pages数据的文档,示例:
确保仅当"/document/pages/*?exists(/document/pages/*)"/document/pages/*存在时才尝试生成投影。 - 修正技能执行顺序并添加空值处理:将SplitSkill移至EmbeddingSkill之前(先拆分文档再生成嵌入向量,符合逻辑),同时在投影映射的每个字段中添加
defaultValue,避免因源字段为空导致投影失败,示例:{ "name": "chunk", "source": "/document/pages/*", "inputs": [{"name": "value", "source": "/document/pages/*", "defaultValue": ""}] } - 启用删除操作跳过技能执行:在索引器配置的
parameters中设置skipSkillExecutionForDeletes: true,针对删除操作直接跳过技能集处理,仅执行索引删除,不触发投影逻辑。
内容的提问来源于stack exchange,提问作者Hessel
相关产品推荐
相关产品推荐

