Azure Document Intelligence自定义模型未生效问题求助
问题描述
我是编程新手,在Document Intelligence Studio里创建了自定义提取模板模型,添加自定义标签后训练、测试效果都正常。但用Python代码调用该模型ID分析PDF文件时,模型完全没应用自定义标签,反而像是在用自动标注功能——这不是我想要的,不然我直接用示例模型就行了。尝试更换Studio里的模型类型也没用,怀疑是代码问题(代码基本是Studio提供的示例),附上代码求排查:
endpoint = "<myendpoint>" key = "<mykey>" model_id = "<my-custom-model-id>" formUrl = "<pdf-url>" document_analysis_client = DocumentAnalysisClient( endpoint=endpoint, credential=AzureKeyCredential(key) ) # Initialize the DocumentAnalysisClient with endpoint and key document_analysis_client = DocumentAnalysisClient( endpoint=endpoint, credential=AzureKeyCredential(key) ) # Make sure your document's type is included in the list of document types the custom model can analyze poller = document_analysis_client.begin_analyze_document_from_url(model_id, formUrl) result = poller.result() for idx, document in enumerate(result.documents): print("--------Analyzing document #{}--------".format(idx + 1)) print("Document has type {}".format(document.doc_type)) print("Document has confidence {}".format(document.confidence)) print("Document was analyzed by model with ID {}".format(result.model_id)) for name, field in document.fields.items(): field_value = field.value if field.value else field.content print("......found field of type '{}' with value '{}' and with confidence {}".format(field.value_type, field_value, field.confidence)) # iterate over tables, lines, and selection marks on each page for page in result.pages: print("\nLines found on page {}".format(page.page_number)) for line in page.lines: print("...Line '{}'".format(line.content.encode('utf-8'))) for word in page.words: print( "...Word '{}' has a confidence of {}".format( word.content.encode('utf-8'), word.confidence ) ) for selection_mark in page.selection_marks: print( "...Selection mark is '{}' and has a confidence of {}".format( selection_mark.state, selection_mark.confidence ) ) for i, table in enumerate(result.tables): print("\nTable {} can be found on page:".format(i + 1)) for region in table.bounding_regions: print("...{}".format(i + 1, region.page_number)) for cell in table.cells: print( "...Cell[{}][{}] has content '{}'".format( cell.row_index, cell.column_index, cell.content.encode('utf-8') ) ) print("-----------------------------------")
我是否需要在代码中指定标签?或者还有其他问题?
排查方案
核心问题分析
你不需要在代码里手动指定标签,自定义模板模型的标签是训练时固化好的,API调用会自动返回对应字段。出现这个问题大概率是以下几个原因:
- 模型未发布:训练完成不代表模型能被API调用,必须在Document Intelligence Studio里完成发布操作,获取发布后的正式模型ID。
- 模型ID错误:代码里的
model_id可能是训练过程中的临时ID,而非发布后的正式ID,去Studio的模型发布页面复制正确的ID。 - 客户端重复初始化(冗余但无功能影响):你的代码里重复初始化了
document_analysis_client,可以删掉重复的那部分。
修正后的代码
保留核心逻辑,去掉冗余部分,重点关注自定义字段的输出:
endpoint = "<myendpoint>" key = "<mykey>" # 替换成发布后的正式模型ID model_id = "<published-custom-template-model-id>" formUrl = "<pdf-url>" # 仅初始化一次客户端 document_analysis_client = DocumentAnalysisClient( endpoint=endpoint, credential=AzureKeyCredential(key) ) # 调用自定义模型分析文档 poller = document_analysis_client.begin_analyze_document_from_url(model_id, formUrl) result = poller.result() # 重点查看自定义标签对应的字段 for idx, document in enumerate(result.documents): print(f"--------分析文档 #{idx + 1}--------") print(f"文档类型: {document.doc_type}") print(f"使用的模型ID: {result.model_id}") for name, field in document.fields.items(): field_value = field.value if field.value else field.content print(f"......找到自定义字段 '{name}',类型: {field.value_type},值: '{field_value}',置信度: {field.confidence}")
额外检查步骤
- 确认Studio里的模型状态是已发布,且发布时选择的是训练成功的版本。
- 测试用的PDF文件要和训练时的文件格式、布局一致,避免因格式差异导致模型识别失效。
- 检查API密钥和端点是否对应正确的Azure资源(比如不要混用不同区域的资源)。
内容的提问来源于stack exchange,提问作者HaniaZ8
相关产品推荐
相关产品推荐

