Google Document AI自定义抽取器Python代码报错:Property无OccurenceType属性
Google Document AI 自定义抽取器代码报错解决
问题概述
已训练好的Google Document AI Custom Extractor模型在云控制台测试正常,但编写Python代码时遇到两个问题:
- 添加自定义字段实体定义(如
cancerType)后,报错:An error occurred: type object 'Property' has no attribute 'OccurenceType' - 不定义实体时,触发400错误,提示缺少schema
自定义字段配置:
| 名称 | 数据类型 | 出现次数 | 是否启用 |
|---|---|---|---|
| cancerType | Plain text | Optional once | Yes |
| mutatedGene | Plain text | Optional multiple | Yes |
| otherBiomarker | Plain text | Optional multiple | Yes |
| patientDOB | Plain text | Optional once | Yes |
| patientGender | Plain text | Optional once | Yes |
| patientName | Plain text | Optional once | Yes |
| reportDate | Plain text | Optional once | Yes |
| specimenDate | Plain text | Optional once | Yes |
报错原因
- 枚举名称拼写错误:正确的枚举名称是
OccurrenceType(原代码少了一个'r'),且该枚举属于documentai.DocumentSchema.EntityType,并非Property类的属性。 - 无需手动指定schema_override:已训练完成的Custom Extractor模型,其schema已绑定到对应的处理器版本,控制台测试正常说明模型配置正确,代码中手动添加
schema_override反而会引发冲突或错误。
修正后的代码
from typing import Optional from google.api_core.client_options import ClientOptions from google.cloud import documentai_v1beta3 as documentai project_id = "xxx" location = "xxxx" processor_id = "xxxx" file_path = "report_f1.pdf" mime_type = "application/pdf" processor_version_id = "xxxx" def process_document_custom_extractor_sample( project_id: str, location: str, processor_id: str, processor_version: str, file_path: str, mime_type: str, ) -> None: try: # 已训练的自定义抽取器无需手动定义schema,直接调用接口即可 document = process_document( project_id, location, processor_id, processor_version, file_path, mime_type, ) for entity in document.entities: print_entity(entity) # 打印嵌套实体(如果存在) for prop in entity.properties: print_entity(prop) except Exception as e: print(f"发生错误: {e}") def print_entity(entity: documentai.Document.Entity) -> None: key = entity.type_ text_value = entity.text_anchor.content confidence = entity.confidence normalized_value = entity.normalized_value.text if entity.normalized_value else None print(f" * {repr(key)}: {repr(text_value)}({confidence:.1%} 置信度)") if normalized_value: print(f" * 标准化值: {repr(normalized_value)}") def process_document( project_id: str, location: str, processor_id: str, processor_version: str, file_path: str, mime_type: str, ) -> documentai.Document: try: # 非"us"区域需指定对应api_endpoint client = documentai.DocumentProcessorServiceClient( client_options=ClientOptions(api_endpoint=f"{location}-documentai.googleapis.com") ) # 构造处理器版本的完整资源名称 name = client.processor_version_path( project_id, location, processor_id, processor_version ) # 读取目标文件内容 with open(file_path, "rb") as image: image_content = image.read() # 配置处理请求 request = documentai.ProcessRequest( name=name, raw_document=documentai.RawDocument(content=image_content, mime_type=mime_type), ) result = client.process_document(request=request) return result.document except Exception as e: raise RuntimeError(f"处理文档错误: {e}") # 调用核心处理函数 process_document_custom_extractor_sample(project_id, location, processor_id, processor_version_id, file_path, mime_type)
关键修正说明
- 移除了手动定义
schema_override的代码块,避免与模型绑定的schema冲突 - 优化了
print_entity函数的空值判断逻辑,防止标准化值为空时引发报错 - 精简了请求参数,保留自定义抽取器所需的核心配置
内容的提问来源于stack exchange,提问作者Matt Reidy
相关产品推荐
相关产品推荐

