You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Azure AI Search索引投影提取PDF单页内容遇报错求助

问题描述

需使用Azure AI Search返回PDF文档中匹配搜索查询的特定页面。当前通过generateNormalizedImagePerPage将PDF每页转为图片,再用OcrSkill提取文本,虽能拆分内容,但查询索引时返回的是整个PDF文档,而非匹配的单页。

尝试通过索引投影将PDF每页作为搜索索引中的独立文档,创建索引及投影配置后,创建技能集时触发报错:

One or more index projection selectors are invalid. 
Details: Index 'myindex' must contain field 'content', it must be of type Edm.String, 
cannot be the key field and it must be filterable.

用户的索引定义代码:

var index = new SearchIndex(name: "myindex")
{
    Fields =
    [
        new SearchField (name: "id", type: SearchFieldDataType.String) 
            { IsSearchable = true, IsKey = true, },
        new SearchField (name: "content", type: SearchFieldDataType.String) 
            { IsFilterable = true, IsKey = false },
        new SearchField (name: "pagetext", type: SearchFieldDataType.String) 
            { IsSearchable = true },
        new SearchField (name: "pagenumber", type: SearchFieldDataType.String) 
            { IsSearchable = true }
    ]
};

用户的索引投影配置代码:

var mappings = new List<InputFieldMappingEntry>
{
    new (name: "pagetext")
    {
        Source = "/document/normalized_images/*/text"
    },
    new (name: "pagenumber")
    {
        Source = "/document/normalized_images/*/pageNumber"
    }
};

var selectors = new List<SearchIndexerIndexProjectionSelector>
{
    new (targetIndexName: "myindex",
         parentKeyFieldName: "content",
         sourceContext: "/document/normalized_images/*",
         mappings: mappings)
};

var indexProjections = new SearchIndexerIndexProjections(selectors)
{
    Parameters = new SearchIndexerIndexProjectionsParameters
    {
        ProjectionMode = IndexProjectionMode.SkipIndexingParentDocuments
    }
};
解决方法

报错核心原因是:索引投影的parentKeyFieldName指定了content字段,但投影配置中未将父文档的对应字段映射到子文档的content字段,导致系统检测不到该字段的有效赋值逻辑,触发校验报错。

具体修正步骤:

  1. 补充content字段的投影映射
    在投影字段映射列表中添加content字段的映射,将父文档的content字段(原始PDF的标识字段)传递给子文档:
    var mappings = new List<InputFieldMappingEntry>
    {
        new (name: "pagetext")
        {
            Source = "/document/normalized_images/*/text"
        },
        new (name: "pagenumber")
        {
            Source = "/document/normalized_images/*/pageNumber"
        },
        // 新增:将父文档的content字段映射到子文档的content字段
        new (name: "content")
        {
            Source = "/document/content"
        }
    };
    
  2. 确保父文档content字段可被提取
    确认索引器配置中已开启原始PDF的文本提取功能,保证父文档的content字段能正常获取到值(可通过内置的PDF解析能力实现)。
  3. 优化pagenumber字段类型(可选)
    将pagenumber字段类型改为Int32,更便于后续筛选、排序操作,修改后的索引字段定义:
    new SearchField (name: "pagenumber", type: SearchFieldDataType.Int32) 
        { IsFilterable = true, IsSortable = true }
    

内容的提问来源于stack exchange,提问作者Tyloo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 22:40:53