You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用Google AutoML实体提取模型处理PDF报404错误排查

问题描述
  • 目标:基于已训练完成的Google AutoML Natural Language实体提取模型,从PDF文档中提取实体及关联值,最终导出为CSV文件,当前处于Python调用代码的开发调试阶段。
  • 现象:参考AutoML UI提供的示例代码编写预测逻辑,传入已上传到Cloud Storage的PDF文件路径调用预测接口时,持续抛出404相关报错,无法定位根因。
测试代码
import sys
import os

from google.api_core.client_options import ClientOptions
from google.cloud import automl_v1

os.environ["GOOGLE_APPLICATION_CREDENTIALS"]="/Users/jetsonwu/intelligent-upload/AutoML_NLP/GCP/automl-pipeline-354119-bc09422edc36.json"

model_id = "TEN7692325184920879104"
file_path = "gs://intelligent_upload/electric-bill/pdf/China World Trade Center EB1.pdf"

def inline_text_payload(file_path):
    with open(file_path, 'rb') as ff:
        content = ff.read()
    return {'text_snippet': {'content': content, 'mime_type': 'text/plain'} }

def pdf_payload(file_path):
    return {'document': {'input_config': {'gcs_source': {'input_uris': [file_path] } } } }

def get_prediction(file_path, model_name):
    options = ClientOptions(api_endpoint='us-automl.googleapis.com')
    prediction_client = automl_v1.PredictionServiceClient(client_options=options)

    # payload = inline_text_payload(file_path)
    # Uncomment the following line (and comment the above line) if want to predict on PDFs.
    payload = pdf_payload(file_path)
    params = {}
    request = prediction_client.predict(name=model_name, payload=payload, params=params)
    return request  # waits until request is returned

if __name__ == '__main__':
    file_path = sys.argv[1]
    model_name = sys.argv[2]

    print(get_prediction(file_path, model_id))
报错信息
E0627 22:10:44.828979000 4336043392 hpack_parser.cc:1234]              Error parsing metadata: error=invalid value key=content-type value=text/html; charset=UTF-8

---------------------------------------------------------------------------
_InactiveRpcError                         Traceback (most recent call last)
File ~/Library/Python/3.8/lib/python/site-packages/google/api_core/grpc_helpers.py:50, in _wrap_unary_errors.<locals>.error_remapped_callable(*args, **kwargs)
     49 try:
---&gt; 50     return callable_(*args, **kwargs)
     51 except grpc.RpcError as exc:

File ~/Library/Python/3.8/lib/python/site-packages/grpc/_channel.py:946, in _UnaryUnaryMultiCallable.__call__(self, request, timeout, metadata, credentials, wait_for_ready, compression)
    944 state, call, = self._blocking(request, timeout, metadata, credentials,
    945                               wait_for_ready, compression)
--&gt; 946 return _end_unary_response_blocking(state, call, False, None)

File ~/Library/Python/3.8/lib/python/site-packages/grpc/_channel.py:849, in _end_unary_response_blocking(state, call, with_call, deadline)
    848 else:
--&gt; 849     raise _InactiveRpcError(state)

_InactiveRpcError: <_InactiveRpcError of RPC that terminated with:
    status = StatusCode.UNIMPLEMENTED
    details = "Received http2 header with status: 404"
    debug_error_string = "{"created":"@1656382244.829029000","description":"Error received from peer ipv4:142.250.191.234:443","file":"src/core/lib/surface/call.cc","file_line":967,"grpc_message":"Received http2 header with status: 404","grpc_status":12}"
>

The above exception was the direct cause of the following exception:

MethodNotImplemented                      Traceback (most recent call last)
...
     50     return callable_(*args, **kwargs)
     51 except grpc.RpcError as exc:
---&gt; 52     raise exceptions.from_grpc_error(exc) from exc

MethodNotImplemented: 501 Received http2 header with status: 404
问题根因及修复方案

报错核心是请求命中了不存在的接口路径,服务端返回404状态的HTML页面,导致gRPC解析元数据失败,具体问题点有三个:

  • 模型资源名格式错误
    接口要求传入的name参数不能只填model_id,必须是完整的资源路径,格式为projects/{GCP项目ID}/locations/{模型部署区域}/models/{model_id},结合当前配置,正确的模型名应为projects/automl-pipeline-354119/locations/us/models/TEN7692325184920879104。另外main函数存在参数传递bug:定义了接收sys.argv[2]作为model_name,实际调用时却硬编码传入了model_id变量。
  • 在线predict接口不支持PDF直接输入
    代码中使用的在线同步预测predict接口,仅支持传入长度1万字符以内的纯文本片段,入参结构只识别text_snippet字段。代码中构造的pdf_payload传入了document结构的GCS文件路径,完全不符合该接口的入参规范,直接触发接口路由404。UI里复制的代码是通用模板,注释里标注的PDF预测支持是错误的,不适用于AutoML Natural Language的在线预测接口。
  • PDF预测需选对接口类型
    要处理PDF文件有两种可选方案:
    • 实时预测场景:先通过PDF解析工具提取文件内的纯文本内容,将文本转成utf-8编码的字符串后,用inline_text_payload格式传入predict接口调用,注意修正原inline_text_payload函数直接读二进制bytes未解码的问题。
    • 批量处理场景:改用异步批量预测batch_predict接口,该接口支持直接传入GCS路径下的PDF作为输入,预测完成后结果会自动写入指定的GCS输出路径,适合大批量文件处理,调用后需要等待异步任务执行完成再读取结果。

内容的提问来源于stack exchange,提问作者Jetson Earth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 12:06:51