You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Vertex AI批量预测中的‘Failed to load model’错误?

问题描述

我正尝试在Vertex AI中为自定义训练的模型运行批量预测。该模型以scikit-learn框架的model.joblib文件形式存储,从Cloud Storage桶导入至模型注册表。

为了调度批量预测,我编写了一个Cloud Function实现该功能,额外添加了以下资源配置:

"dedicated_resources": {
        "machine_spec": {
            "machine_type": "c2-standard-16"
        },
        "starting_replica_count":1,
        "max_replica_count":3,
}

我需要为约30万用户、47个特征的7聚类KMeans模型进行预测,但持续遇到以下错误:

  1. 日志报错:

ERROR:root:Failed to load model: Could not load the model: /tmp/model/0001/model.joblib. std::bad_alloc. (Error code: 0)

  1. 批量预测UI报错:

Error: model server never became ready. Please validate that your model file or container configuration are valid.

我已尝试多种机器类型:

  • c2-standard-16
  • n1-highcpu-4
  • n1-highcpu-16

但错误仍未改变。请问这是机器类型问题还是其他原因?该如何解决?

附批量预测的Cloud Function代码:

from google.cloud import aiplatform_v1beta1
from google.protobuf import json_format
from google.protobuf.struct_pb2 import Value

def hello_pubsub(event, context):
  project = "abc"
  display_name = "audience-segmentation-predictions"
  model_name = "projects/abc/locations/us-central1/models/19290"
  instances_format = "bigquery"
  bigquery_source_input_uri = "bq://abc.ml_final_g0_us.model_input_transformed_daily"
  predictions_format = "bigquery"
  bigquery_destination_output_uri = "bq://abc.ml_final_g0_us"
  location = "us-central1"
  api_endpoint = "us-central1-aiplatform.googleapis.com"


  def create_batch_prediction_job_bigquery_sample(project,display_name,model_name,instances_format,bigquery_source_input_uri,predictions_format,bigquery_destination_output_uri,location,api_endpoint):
      # The AI Platform services require regional API endpoints.
      client_options = {"api_endpoint": api_endpoint}
      # Initialize client that will be used to create and send requests.
      # This client only needs to be created once, and can be reused for multiple requests.
      client = aiplatform_v1beta1.JobServiceClient(client_options=client_options)
      model_parameters_dict = {}
      model_parameters = json_format.ParseDict(model_parameters_dict, Value())

      batch_prediction_job = {
          "display_name": display_name,
          # Format: 'projects/{project}/locations/{location}/models/{model_id}'
          "model": model_name,
          "model_parameters": model_parameters,
          "input_config": {
              "instances_format": instances_format,
              "bigquery_source": {"input_uri": bigquery_source_input_uri},
          },
          "output_config": {
              "predictions_format": predictions_format,
              "bigquery_destination": {"output_uri": bigquery_destination_output_uri},
          },
          "instance_config": {"excluded_fields": "clientid"},
          "dedicated_resources": {
            "machine_spec": {
                "machine_type": "c2-standard-30"
            },
            
        },
      }
      parent = f"projects/{project}/locations/{location}"
      response = client.create_batch_prediction_job(
          parent=parent, batch_prediction_job=batch_prediction_job
      )
      print("response:", response)

  create_batch_prediction_job_bigquery_sample(project,display_name,model_name,instances_format,bigquery_source_input_uri,predictions_format,bigquery_destination_output_uri,location,api_endpoint)
解决方案

std::bad_alloc错误本质是内存不足,虽然你尝试了不同CPU配置的机器,但可能忽略了内存资源的匹配,或者模型本身存在加载问题,以下是具体排查和解决步骤:

1. 优先调整内存资源配置

你之前选的机器类型内存配比偏低:

  • n1-highcpu-4:仅3.6GB内存
  • n1-highcpu-16:仅14.4GB内存
  • c2-standard-16:64GB内存,但对较大的KMeans模型可能仍不够

建议:

  • 改用高内存机型,比如n1-highmem-16(16vCPU + 104GB内存)或c2-standard-60(60vCPU + 240GB内存),优先保证内存容量足够加载模型。
  • 在dedicated_resources中明确配置高内存机型,避免仅关注CPU核心数。

2. 验证模型文件完整性

  • 本地尝试加载model.joblib文件,确认模型本身无损坏:
    import joblib
    model = joblib.load("model.joblib")
    # 用少量测试数据运行预测,验证模型功能正常
    
    如果本地加载也报错,说明模型文件在导出或上传时损坏,需要重新训练并导出模型。
  • 查看model.joblib的文件大小,如果超过几十GB,检查训练时是否存储了不必要的大对象(比如训练数据集副本),重新导出时只保留模型必要参数。

3. 拆分批量预测任务

将BigQuery中的30万条输入数据拆分为多个小批量任务运行,减少单任务的内存占用压力,比如按日期或用户ID分段处理。

4. 检查容器版本兼容性

Vertex AI运行scikit-learn模型使用官方容器,若训练时的scikit-learn版本与容器版本不一致,可能导致模型加载失败:

  • 确认训练环境的scikit-learn版本,与Vertex AI批量预测容器的版本匹配。
  • 若版本不兼容,可自定义容器镜像,使用和训练环境一致的依赖版本,再导入模型注册表进行批量预测。

内容的提问来源于stack exchange,提问作者Avantika Banerjee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 23:47:41