Azure AI向量搜索:如何基于父元数据过滤分块PDF
解决Azure搜索分块数据的父元数据映射问题
核心问题分析
你通过Azure门户的「Import and vectorize」快速配置向量搜索后,分块数据无法继承父文档的prosjektnummer元数据,导致无法基于该字段过滤分块,且索引器运行时出现0/0无操作的情况。主要原因有两点:
- Blob自定义元数据在Azure搜索中会自动添加
metadata_前缀,你之前的字段映射未使用正确的源字段名 - 技能集的索引投影配置中,未将父元数据映射到分块索引字段
最简解决方案
1. 修正索引器的字段映射
修改索引器的fieldMappings,将prosjektnummer的源字段名改为metadata_prosjektnummer(Blob自定义元数据的标准命名格式):
"fieldMappings": [ { "sourceFieldName": "metadata_storage_name", "targetFieldName": "title", "mappingFunction": null }, { "sourceFieldName": "metadata_prosjektnummer", "targetFieldName": "prosjektnummer", "mappingFunction": null } ]
2. 更新技能集的索引投影配置
在技能集的indexProjections.selectors.mappings中添加prosjektnummer的映射,确保父文档的元数据被传递到每个分块:
"mappings": [ { "name": "chunk", "source": "/document/pages/*", "sourceContext": null, "inputs": [] }, { "name": "vector", "source": "/document/pages/*/vector", "sourceContext": null, "inputs": [] }, { "name": "title", "source": "/document/metadata_storage_name", "sourceContext": null, "inputs": [] }, { "name": "prosjektnummer", "source": "/document/prosjektnummer", "sourceContext": null, "inputs": [] } ]
3. 重置并重新运行索引器
索引器出现0/0无操作是因为增量索引机制认为文档未变更,需先重置索引器状态再全量运行:
- 在Azure门户中找到目标索引器,点击「重置」按钮
- 重置完成后,点击「运行」按钮触发全量索引
可选方案说明
若考虑直接过滤父文档,可修改技能集的projectionMode为includeParentDocuments,此时索引中将同时存储父文档和分块数据:
- 父文档包含完整元数据,可先过滤父文档得到
parent_id - 再基于
parent_id过滤对应的分块数据
但该方案会增加索引存储量,对于你的场景,将元数据直接映射到分块是更简洁的选择
内容的提问来源于stack exchange,提问作者Trygve Leithe Svalheim
相关产品推荐
相关产品推荐

