You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何通过Python调用训练后的Azure Form Recognizer提取自定义PDF数据

问题:使用Azure Form Recognizer自定义模型提取PDF数据失败

我需要用Python调用训练好的Azure Form Recognizer自定义模型,提取自定义PDF中的数据。已获取模型生成的代码,参数配置正确,但始终无法成功读取数据。(已查阅所有相关StackOverflow问题,未找到适配我场景的解决方案)

官方生成的代码示例

"""
This code sample shows Custom Extraction Model operations with the Azure Form Recognizer client library. 
The async versions of the samples require Python 3.6 or later.
"""

from azure.core.credentials import AzureKeyCredential
from azure.ai.formrecognizer import DocumentAnalysisClient

"""
Remember to remove the key from your code when you're done, and never post it publicly. For production, use
secure methods to store and access your credentials.
"""
endpoint = "YOUR_FORM_RECOGNIZER_ENDPOINT"
key = "YOUR_FORM_RECOGNIZER_KEY"

model_id = "YOUR_CUSTOM_BUILT_MODEL_ID"
formUrl = "YOUR_DOCUMENT"

document_analysis_client = DocumentAnalysisClient(
    endpoint=endpoint, credential=AzureKeyCredential(key)
)

# Make sure your document's type is included in the list of document types the custom model can analyze
poller = document_analysis_client.begin_analyze_document_from_url(model_id, formUrl)
result = poller.result()

for idx, document in enumerate(result.documents):
    print("--------Analyzing document #{}--------".format(idx + 1))
    print("Document has type {}".format(document.doc_type))
    print("Document has confidence {}".format(document.confidence))
    print("Document was analyzed by model with ID {}".format(result.model_id))
    for name, field in document.fields.items():
        field_value = field.value if field.value else field.content
        print("......found field of type '{}' with value '{}' and with confidence {}".format(field.value_type, field_value, field.confidence))


# iterate over tables, lines, and selection marks on each page
for page in result.pages:
    print("\nLines found on page {}".format(page.page_number))
    for line in page.lines:
        print("...Line '{}'".format(line.content.encode('utf-8')))
    for word in page.words:
        print(
            "...Word '{}' has a confidence of {}".format(
                word.content.encode('utf-8'), word.confidence
            )
        )
    for selection_mark in page.selection_marks:
        print(
            "...Selection mark is '{}' and has a confidence of {}".format(
                selection_mark.state, selection_mark.confidence
            )
        )

for i, table in enumerate(result.tables):
    print("\nTable {} can be found on page:".format(i + 1))
    for region in table.bounding_regions:
        print("...{}".format(region.page_number))
    for cell in table.cells:
        print(
            "...Cell[{}][{}] has content '{}'".format(
                cell.row_index, cell.column_index, cell.content.encode('utf-8')
            )
        )
print("-----------------------------------")

我读取Blob内容的方式

读取Blob内容的方式

排查与解决步骤

  • 确认Blob访问权限:确保生成的Blob URL包含有效的SAS令牌,令牌需具备读取权限且未过期。私有Blob的原始URL无法被Form Recognizer访问。
  • 更换文档分析方法:若直接读取Blob字节流,不要使用begin_analyze_document_from_url,改用begin_analyze_document传入字节数据,示例代码:
# 假设已通过Azure Blob客户端获取到blob_client对象
blob_data = blob_client.download_blob().readall()

# 替换原有的URL分析逻辑
poller = document_analysis_client.begin_analyze_document(model_id, blob_data)
result = poller.result()
  • 验证模型与PDF布局一致性:自定义模型依赖训练时的文档布局,若目标PDF与训练样本布局存在差异(如字段位置、格式变化),会导致提取失败。需确保待分析PDF与训练样本布局完全匹配。
  • 添加异常捕获排查错误:在代码中加入异常处理,获取具体错误信息,定位问题根源:
try:
    poller = document_analysis_client.begin_analyze_document_from_url(model_id, formUrl)
    result = poller.result()
except Exception as e:
    print(f"分析失败: {str(e)}")
  • 更新SDK版本:确保azure-ai-formrecognizer SDK为最新版本,执行以下命令更新:
pip install --upgrade azure-ai-formrecognizer

内容的提问来源于stack exchange,提问作者WhoamI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 10:52:51