如何在Azure Document Intelligence的Markdown结果中保留原文档分页?
拆分Azure Document Intelligence生成的Markdown并保留分页结构
以下是几种可行的解决方案,针对无页码文档的分页拆分问题,同时兼容表格、图片等复杂元素场景:
方法1:结合AnalyzeResult的页面元素与Markdown文本对齐拆分
核心思路是利用服务返回的AnalyzeResult中按页划分的元素(文本行、表格、图片),定位这些元素在完整Markdown中的位置,以此作为分页边界,避免直接用lines重建文本的格式差异问题。
具体步骤
- 先获取完整的Markdown文本(
result.content) - 遍历每个页面,收集该页所有元素的特征文本(文本行内容、表格首尾行、图片alt文本)
- 从后往前匹配当前页最后一个元素在Markdown中的位置,结合元素类型(表格/图片)调整拆分边界,确保不破坏结构化内容
- 按拆分点切割Markdown,添加分页标识
代码实现
from azure.ai.documentintelligence import DocumentIntelligenceClient from azure.ai.documentintelligence.models import ContentFormat, AnalyzeResult import os def split_markdown_by_pages(result: AnalyzeResult): full_markdown = result.content page_splits = [] current_pos = 0 for page_num, page in enumerate(result.pages, start=1): page_elements = [] # 收集当前页的文本行 for line in page.lines: cleaned_content = line.content.strip() if cleaned_content: page_elements.append(cleaned_content) # 收集当前页的表格首尾行文本 for table in result.tables: if table.bounding_regions[0].page_number == page_num: table_lines = [line.content.strip() for line in result.lines if any(reg.page_number == page_num for reg in line.bounding_regions) and line.content in table.content] if table_lines: page_elements.extend([table_lines[0], table_lines[-1]]) # 收集当前页的图片alt文本 for image in result.images: if image.bounding_regions[0].page_number == page_num and image.content: page_elements.append(image.content.strip()) # 定位当前页最后一个可匹配元素的位置 last_match_pos = current_pos for elem in reversed(page_elements): elem_start = full_markdown.find(elem, current_pos) if elem_start != -1: last_match_pos = elem_start + len(elem) break # 针对表格/图片调整拆分边界,避免拆分结构化内容 # 表格:找到表格后的第一个空行 has_table = any(t.bounding_regions[0].page_number == page_num for t in result.tables) if has_table: table_end = full_markdown.find('\n\n', last_match_pos) if table_end != -1: last_match_pos = table_end + 2 # 图片:找到图片标记的结束括号 has_image = any(i.bounding_regions[0].page_number == page_num for i in result.images) if has_image: img_end = full_markdown.find(')', last_match_pos) if img_end != -1: last_match_pos = img_end + 1 # 截取当前页Markdown并添加分页标识 page_content = full_markdown[current_pos:last_match_pos].strip() page_splits.append(f"<!-- Page {page_num} -->\n{page_content}") current_pos = last_match_pos # 处理剩余未拆分的内容 if current_pos < len(full_markdown): page_splits.append(f"<!-- Page {len(result.pages)+1} -->\n{full_markdown[current_pos:].strip()}") return page_splits # 使用示例 client = DocumentIntelligenceClient( endpoint=os.getenv("DOCUMENT_INTELLIGENCE_ENDPOINT"), credential=AzureKeyCredential(os.getenv("DOCUMENT_INTELLIGENCE_KEY")) ) with open("target-document.pdf", "rb") as f: poller = client.begin_analyze_document( "prebuilt-layout", bytes_source=f.read(), output_content_format=ContentFormat.MARKDOWN ) result = poller.result() # 拆分并输出分页后的Markdown for idx, page_md in enumerate(split_markdown_by_pages(result), 1): print(f"=== 第{idx}页 ===\n{page_md}\n")
方法2:利用页面布局的位置信息辅助拆分
对于无页码的文档,可以通过元素在页面中的坐标位置判断分页边界:
- 定义页面底部区域(比如页面高度的85%以上)为页脚区域
- 遍历每个页面的所有元素,找到最后一个不在页脚区域内的元素,以此作为分页结束点
- 结合该元素在Markdown中的位置完成拆分
这种方法适合布局规范的文档,能避免文本匹配的误差。
注意事项
- 对于跨页表格,需要额外判断表格的跨页属性(
table.spans),避免拆分跨页表格内容 - 若文档存在大量无意义空白行,可先对Markdown文本做预处理(合并连续空白行),提升匹配准确性
- 不同文档的格式差异可能需要调整元素匹配、边界判断的逻辑,建议根据实际场景优化
内容的提问来源于stack exchange,提问作者filippocurati
相关产品推荐
相关产品推荐

