You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Azure搜索索引字段添加文档分块所属的原页码?

给Azure集成向量化的文档分块添加PDF页码的实现方案

Azure Integrated Vectorization处理PDF时会自动捕获页码元数据,只要配置得当,就能把页码关联到每个分块,具体用C#实现的步骤如下:

1. 确认数据源元数据提取

默认情况下,Azure搜索处理可提取文本的PDF(非扫描件)时,会自动把页码提取到内置元字段metadata_storage_pdf_page_number,无需额外配置数据源。

2. 给索引新增页码字段

修改索引定义,添加一个用于存储页码的字段,示例代码:

var index = new SearchIndex("your-index-name")
{
    Fields =
    {
        new SimpleField("id", SearchFieldDataType.String) { IsKey = true },
        new SearchableField("content") { IsFilterable = false, IsSortable = false },
        new VectorField("content_vector", SearchFieldDataType.Collection(SearchFieldDataType.Single))
        {
            Dimensions = 1536,
            VectorSearchConfigurationName = "your-vector-config"
        },
        // 新增页码字段,支持筛选和排序
        new SimpleField("pdf_page_number", SearchFieldDataType.Int32) { IsFilterable = true, IsSortable = true }
    }
};

3. 字段映射关联页码

在索引器中,把内置的页码元字段映射到你新增的索引字段,示例:

var indexer = new SearchIndexer("your-indexer-name", "your-data-source", "your-index")
{
    FieldMappings =
    {
        // 将PDF页码映射到索引字段
        new FieldMapping("metadata_storage_pdf_page_number", "pdf_page_number"),
        // 其他必要的字段映射
        new FieldMapping("content", "content"),
        new FieldMapping("id", "id")
    },
    SkillsetName = "your-skillset-name"
};

4. 检索时用页码实现跳转

索引完成后,每个文档分块都会关联对应的PDF页码。展示引用内容时,直接获取pdf_page_number的值,给PDF链接加上锚点#page=[页码],即可引导用户跳转到PDF的对应页面。

额外提醒

  • 若处理的是扫描版PDF(图片格式),Azure无法直接提取页码,需先配置OCR技能完成文本识别,同时在OCR技能中开启页码提取。
  • 若使用自定义文档拆分逻辑,需在拆分过程中手动传递页码信息,可通过自定义技能实现。

内容的提问来源于stack exchange,提问作者Punit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 16:28:16