You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Azure Document Intelligence获取按页分隔的Markdown内容?

解决Azure Document Intelligence按页获取Markdown内容的问题

方案1:通过段落页码信息手动拆分完整Markdown

prebuilt-layout模型返回的结果中,每个段落(paragraphs)都附带page_number属性。你可以先获取完整的Markdown内容,再结合段落的页码标记,将内容按页拆分。

示例代码:

from azure.ai.documentintelligence import DocumentIntelligenceClient
from azure.core.credentials import AzureKeyCredential
from azure.ai.documentintelligence.models import ContentFormat, AnalyzeDocumentRequest

# 初始化客户端
document_intelligence_client = DocumentIntelligenceClient(endpoint=endpoint, credential=AzureKeyCredential(key))

# 请求全文档分析,获取Markdown和段落信息
poller = document_intelligence_client.begin_analyze_document(
    "prebuilt-layout",
    analyze_request=AnalyzeDocumentRequest(base64_source=doc_bytes),
    output_content_format=ContentFormat.MARKDOWN
)
result = poller.result()

# 按页码分组段落内容
page_content_map = {}
for para in result.paragraphs:
    page_num = para.page_number
    if page_num not in page_content_map:
        page_content_map[page_num] = []
    page_content_map[page_num].append(para.content)

# 生成每页的Markdown(保持段落间的换行格式)
page_markdowns = {page: "\n\n".join(content_list) for page, content_list in page_content_map.items()}

# 输出或保存每页内容
for page_num, markdown in page_markdowns.items():
    print(f"=== 第 {page_num} 页 ===")
    print(markdown)

方案2:拆分文档为单页后逐个分析

指定pages参数后耗时未降低,是因为API仍会加载整个文档做预处理。要提升单页分析效率,建议先将原文档拆分为单页文件(比如PDF拆成独立的单页PDF),再对每个单页文件单独调用API,这样每个请求仅处理单页内容,耗时会明显减少。

示例代码(基于已拆分的单页文件):

import os
from azure.ai.documentintelligence import DocumentIntelligenceClient
from azure.core.credentials import AzureKeyCredential
from azure.ai.documentintelligence.models import ContentFormat, AnalyzeDocumentRequest

client = DocumentIntelligenceClient(endpoint=endpoint, credential=AzureKeyCredential(key))
single_page_dir = "./single_page_documents"

# 遍历所有单页文件
for file_name in os.listdir(single_page_dir):
    if not file_name.endswith((".pdf", ".docx")):
        continue
    file_path = os.path.join(single_page_dir, file_name)
    with open(file_path, "rb") as f:
        doc_bytes = f.read()
    
    # 分析单页文档
    poller = client.begin_analyze_document(
        "prebuilt-layout",
        analyze_request=AnalyzeDocumentRequest(base64_source=doc_bytes),
        output_content_format=ContentFormat.MARKDOWN
    )
    result = poller.result()
    
    # 从文件名提取页码(需根据你的命名规则调整)
    page_num = file_name.split("-")[-1].split(".")[0]
    print(f"=== 第 {page_num} 页 ===")
    print(result.content)

关键说明

目前prebuilt-layout模型的pages字段仅返回结构化元素(如行、单词、表格坐标等),不会直接提供Markdown格式内容。所有Markdown内容统一在返回结果的顶层content字段中,因此必须通过上述两种方式实现分页获取。

内容的提问来源于stack exchange,提问作者Nikola Petrovic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 17:13:20