You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure文档智能同类型表单布局不一致及图片调用问题求助

问题1:Prebuilt-Layout模型返回结构不一致的处理方案

Prebuilt-Layout模型的sections字段是基于文档的视觉区块(如页眉、正文分区、页脚等)自动识别生成的,同类型表单返回结构差异通常源于表单的细微视觉变化(比如空白区域大小、文本排版偏移、扫描件噪声等),导致模型对区块的划分逻辑触发不同。可以通过以下方式统一处理逻辑:

  1. 统一段落提取逻辑:
    不管返回结果是否包含sections,编写通用代码提取所有段落,避免依赖固定结构。示例代码如下:

    def extract_all_paragraphs(result):
        paragraphs = []
        # 优先处理sections内的段落
        if hasattr(result, 'sections') and result.sections:
            for section in result.sections:
                if hasattr(section, 'paragraphs') and section.paragraphs:
                    paragraphs.extend(section.paragraphs)
        # 无sections时直接取根级段落
        else:
            if hasattr(result, 'paragraphs') and result.paragraphs:
                paragraphs.extend(result.paragraphs)
        return paragraphs
    
    # 调用示例
    layout_result = layout(pdf_path)
    all_paragraphs = extract_all_paragraphs(layout_result)
    
  2. 优化输入文档质量:

    • 确保所有表单的生成/扫描规格统一(如分辨率≥300DPI、无倾斜、无多余噪点);
    • 电子生成的PDF避免使用浮动布局或动态内容,尽量采用固定排版的静态PDF。
  3. 切换更适合表单的模型:
    如果核心需求是提取表单的结构化数据(如姓名、出生日期等键值对),建议使用prebuilt-form模型替代prebuilt-layout,该模型专门针对表单场景优化,返回的键值对结构更稳定。

问题2:使用图片格式(JPG/PNG)调用Prebuilt-Layout模型

只需修改请求的content_type为对应图片类型,并读取图片文件的二进制内容即可,模型本身支持图片输入。修改后的示例代码如下:

from azure.ai.documentintelligence import DocumentIntelligenceClient
from azure.ai.documentintelligence.models import ContentFormat

# 假设已初始化document_intelligence_client_sections客户端
def layout_image(image_path):
    # 根据图片后缀设置对应的content_type
    content_type = "image/jpeg" if image_path.lower().endswith(('.jpg', '.jpeg')) else "image/png"
    with open(image_path, "rb") as f:
        poller = document_intelligence_client_sections.begin_analyze_document(
            "prebuilt-layout", 
            f.read(), 
            content_type=content_type,
            output_content_format=ContentFormat.MARKDOWN
        )
    result = poller.result()
    return result

# 调用示例
jpg_result = layout_image("form.jpg")
png_result = layout_image("form.png")

注意:图片输入时需保证画面清晰无模糊,否则会影响布局识别的准确性。

内容的提问来源于stack exchange,提问作者Droid-Bird

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 12:40:08