You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Document AI自定义抽取器Python代码报错:Property无OccurenceType属性

Google Document AI 自定义抽取器代码报错解决

问题概述

已训练好的Google Document AI Custom Extractor模型在云控制台测试正常,但编写Python代码时遇到两个问题:

  • 添加自定义字段实体定义(如cancerType)后,报错:An error occurred: type object 'Property' has no attribute 'OccurenceType'
  • 不定义实体时,触发400错误,提示缺少schema

自定义字段配置:

名称数据类型出现次数是否启用
cancerTypePlain textOptional onceYes
mutatedGenePlain textOptional multipleYes
otherBiomarkerPlain textOptional multipleYes
patientDOBPlain textOptional onceYes
patientGenderPlain textOptional onceYes
patientNamePlain textOptional onceYes
reportDatePlain textOptional onceYes
specimenDatePlain textOptional onceYes

报错原因

  1. 枚举名称拼写错误:正确的枚举名称是OccurrenceType(原代码少了一个'r'),且该枚举属于documentai.DocumentSchema.EntityType,并非Property类的属性。
  2. 无需手动指定schema_override:已训练完成的Custom Extractor模型,其schema已绑定到对应的处理器版本,控制台测试正常说明模型配置正确,代码中手动添加schema_override反而会引发冲突或错误。

修正后的代码

from typing import Optional
from google.api_core.client_options import ClientOptions
from google.cloud import documentai_v1beta3 as documentai

project_id = "xxx"
location = "xxxx"
processor_id = "xxxx"
file_path = "report_f1.pdf"
mime_type = "application/pdf"
processor_version_id = "xxxx"

def process_document_custom_extractor_sample(
    project_id: str,
    location: str,
    processor_id: str,
    processor_version: str,
    file_path: str,
    mime_type: str,
) -> None:
    try:
        # 已训练的自定义抽取器无需手动定义schema,直接调用接口即可
        document = process_document(
            project_id,
            location,
            processor_id,
            processor_version,
            file_path,
            mime_type,
        )

        for entity in document.entities:
            print_entity(entity)
            # 打印嵌套实体(如果存在)
            for prop in entity.properties:
                print_entity(prop)

    except Exception as e:
        print(f"发生错误: {e}")

def print_entity(entity: documentai.Document.Entity) -> None:
    key = entity.type_
    text_value = entity.text_anchor.content
    confidence = entity.confidence
    normalized_value = entity.normalized_value.text if entity.normalized_value else None
    
    print(f"    * {repr(key)}: {repr(text_value)}({confidence:.1%} 置信度)")

    if normalized_value:
        print(f"    * 标准化值: {repr(normalized_value)}")

def process_document(
    project_id: str,
    location: str,
    processor_id: str,
    processor_version: str,
    file_path: str,
    mime_type: str,
) -> documentai.Document:
    try:
        # 非"us"区域需指定对应api_endpoint
        client = documentai.DocumentProcessorServiceClient(
            client_options=ClientOptions(api_endpoint=f"{location}-documentai.googleapis.com")
        )

        # 构造处理器版本的完整资源名称
        name = client.processor_version_path(
            project_id, location, processor_id, processor_version
        )

        # 读取目标文件内容
        with open(file_path, "rb") as image:
            image_content = image.read()

        # 配置处理请求
        request = documentai.ProcessRequest(
            name=name,
            raw_document=documentai.RawDocument(content=image_content, mime_type=mime_type),
        )

        result = client.process_document(request=request)

        return result.document

    except Exception as e:
        raise RuntimeError(f"处理文档错误: {e}")

# 调用核心处理函数
process_document_custom_extractor_sample(project_id, location, processor_id, processor_version_id, file_path, mime_type)

关键修正说明

  • 移除了手动定义schema_override的代码块,避免与模型绑定的schema冲突
  • 优化了print_entity函数的空值判断逻辑,防止标准化值为空时引发报错
  • 精简了请求参数,保留自定义抽取器所需的核心配置

内容的提问来源于stack exchange,提问作者Matt Reidy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 22:57:33