You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Azure Document Intelligence的Markdown结果中保留原文档分页?

拆分Azure Document Intelligence生成的Markdown并保留分页结构

以下是几种可行的解决方案,针对无页码文档的分页拆分问题,同时兼容表格、图片等复杂元素场景:


方法1:结合AnalyzeResult的页面元素与Markdown文本对齐拆分

核心思路是利用服务返回的AnalyzeResult中按页划分的元素(文本行、表格、图片),定位这些元素在完整Markdown中的位置,以此作为分页边界,避免直接用lines重建文本的格式差异问题。

具体步骤

  1. 先获取完整的Markdown文本(result.content)
  2. 遍历每个页面,收集该页所有元素的特征文本(文本行内容、表格首尾行、图片alt文本)
  3. 从后往前匹配当前页最后一个元素在Markdown中的位置,结合元素类型(表格/图片)调整拆分边界,确保不破坏结构化内容
  4. 按拆分点切割Markdown,添加分页标识

代码实现

from azure.ai.documentintelligence import DocumentIntelligenceClient
from azure.ai.documentintelligence.models import ContentFormat, AnalyzeResult
import os

def split_markdown_by_pages(result: AnalyzeResult):
    full_markdown = result.content
    page_splits = []
    current_pos = 0

    for page_num, page in enumerate(result.pages, start=1):
        page_elements = []
        # 收集当前页的文本行
        for line in page.lines:
            cleaned_content = line.content.strip()
            if cleaned_content:
                page_elements.append(cleaned_content)
        # 收集当前页的表格首尾行文本
        for table in result.tables:
            if table.bounding_regions[0].page_number == page_num:
                table_lines = [line.content.strip() for line in result.lines 
                               if any(reg.page_number == page_num for reg in line.bounding_regions) 
                               and line.content in table.content]
                if table_lines:
                    page_elements.extend([table_lines[0], table_lines[-1]])
        # 收集当前页的图片alt文本
        for image in result.images:
            if image.bounding_regions[0].page_number == page_num and image.content:
                page_elements.append(image.content.strip())
        
        # 定位当前页最后一个可匹配元素的位置
        last_match_pos = current_pos
        for elem in reversed(page_elements):
            elem_start = full_markdown.find(elem, current_pos)
            if elem_start != -1:
                last_match_pos = elem_start + len(elem)
                break
        
        # 针对表格/图片调整拆分边界,避免拆分结构化内容
        # 表格:找到表格后的第一个空行
        has_table = any(t.bounding_regions[0].page_number == page_num for t in result.tables)
        if has_table:
            table_end = full_markdown.find('\n\n', last_match_pos)
            if table_end != -1:
                last_match_pos = table_end + 2
        # 图片:找到图片标记的结束括号
        has_image = any(i.bounding_regions[0].page_number == page_num for i in result.images)
        if has_image:
            img_end = full_markdown.find(')', last_match_pos)
            if img_end != -1:
                last_match_pos = img_end + 1
        
        # 截取当前页Markdown并添加分页标识
        page_content = full_markdown[current_pos:last_match_pos].strip()
        page_splits.append(f"<!-- Page {page_num} -->\n{page_content}")
        current_pos = last_match_pos
    
    # 处理剩余未拆分的内容
    if current_pos < len(full_markdown):
        page_splits.append(f"<!-- Page {len(result.pages)+1} -->\n{full_markdown[current_pos:].strip()}")
    
    return page_splits

# 使用示例
client = DocumentIntelligenceClient(
    endpoint=os.getenv("DOCUMENT_INTELLIGENCE_ENDPOINT"),
    credential=AzureKeyCredential(os.getenv("DOCUMENT_INTELLIGENCE_KEY"))
)

with open("target-document.pdf", "rb") as f:
    poller = client.begin_analyze_document(
        "prebuilt-layout",
        bytes_source=f.read(),
        output_content_format=ContentFormat.MARKDOWN
    )
result = poller.result()

# 拆分并输出分页后的Markdown
for idx, page_md in enumerate(split_markdown_by_pages(result), 1):
    print(f"=== 第{idx}页 ===\n{page_md}\n")

方法2:利用页面布局的位置信息辅助拆分

对于无页码的文档,可以通过元素在页面中的坐标位置判断分页边界:

  • 定义页面底部区域(比如页面高度的85%以上)为页脚区域
  • 遍历每个页面的所有元素,找到最后一个不在页脚区域内的元素,以此作为分页结束点
  • 结合该元素在Markdown中的位置完成拆分

这种方法适合布局规范的文档,能避免文本匹配的误差。


注意事项

  • 对于跨页表格,需要额外判断表格的跨页属性(table.spans),避免拆分跨页表格内容
  • 若文档存在大量无意义空白行,可先对Markdown文本做预处理(合并连续空白行),提升匹配准确性
  • 不同文档的格式差异可能需要调整元素匹配、边界判断的逻辑,建议根据实际场景优化

内容的提问来源于stack exchange,提问作者filippocurati

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 18:34:55