使用Azure AI Search索引投影提取PDF单页内容遇报错求助
问题描述
需使用Azure AI Search返回PDF文档中匹配搜索查询的特定页面。当前通过generateNormalizedImagePerPage将PDF每页转为图片,再用OcrSkill提取文本,虽能拆分内容,但查询索引时返回的是整个PDF文档,而非匹配的单页。
尝试通过索引投影将PDF每页作为搜索索引中的独立文档,创建索引及投影配置后,创建技能集时触发报错:
One or more index projection selectors are invalid. Details: Index 'myindex' must contain field 'content', it must be of type Edm.String, cannot be the key field and it must be filterable.
用户的索引定义代码:
var index = new SearchIndex(name: "myindex") { Fields = [ new SearchField (name: "id", type: SearchFieldDataType.String) { IsSearchable = true, IsKey = true, }, new SearchField (name: "content", type: SearchFieldDataType.String) { IsFilterable = true, IsKey = false }, new SearchField (name: "pagetext", type: SearchFieldDataType.String) { IsSearchable = true }, new SearchField (name: "pagenumber", type: SearchFieldDataType.String) { IsSearchable = true } ] };
用户的索引投影配置代码:
var mappings = new List<InputFieldMappingEntry> { new (name: "pagetext") { Source = "/document/normalized_images/*/text" }, new (name: "pagenumber") { Source = "/document/normalized_images/*/pageNumber" } }; var selectors = new List<SearchIndexerIndexProjectionSelector> { new (targetIndexName: "myindex", parentKeyFieldName: "content", sourceContext: "/document/normalized_images/*", mappings: mappings) }; var indexProjections = new SearchIndexerIndexProjections(selectors) { Parameters = new SearchIndexerIndexProjectionsParameters { ProjectionMode = IndexProjectionMode.SkipIndexingParentDocuments } };
解决方法
报错核心原因是:索引投影的parentKeyFieldName指定了content字段,但投影配置中未将父文档的对应字段映射到子文档的content字段,导致系统检测不到该字段的有效赋值逻辑,触发校验报错。
具体修正步骤:
- 补充
content字段的投影映射
在投影字段映射列表中添加content字段的映射,将父文档的content字段(原始PDF的标识字段)传递给子文档:var mappings = new List<InputFieldMappingEntry> { new (name: "pagetext") { Source = "/document/normalized_images/*/text" }, new (name: "pagenumber") { Source = "/document/normalized_images/*/pageNumber" }, // 新增:将父文档的content字段映射到子文档的content字段 new (name: "content") { Source = "/document/content" } }; - 确保父文档
content字段可被提取
确认索引器配置中已开启原始PDF的文本提取功能,保证父文档的content字段能正常获取到值(可通过内置的PDF解析能力实现)。 - 优化
pagenumber字段类型(可选)
将pagenumber字段类型改为Int32,更便于后续筛选、排序操作,修改后的索引字段定义:new SearchField (name: "pagenumber", type: SearchFieldDataType.Int32) { IsFilterable = true, IsSortable = true }
内容的提问来源于stack exchange,提问作者Tyloo
相关产品推荐
相关产品推荐

