如何解决Vertex AI批量预测中的‘Failed to load model’错误?
我正尝试在Vertex AI中为自定义训练的模型运行批量预测。该模型以scikit-learn框架的model.joblib文件形式存储,从Cloud Storage桶导入至模型注册表。
为了调度批量预测,我编写了一个Cloud Function实现该功能,额外添加了以下资源配置:
"dedicated_resources": { "machine_spec": { "machine_type": "c2-standard-16" }, "starting_replica_count":1, "max_replica_count":3, }
我需要为约30万用户、47个特征的7聚类KMeans模型进行预测,但持续遇到以下错误:
- 日志报错:
ERROR:root:Failed to load model: Could not load the model: /tmp/model/0001/model.joblib. std::bad_alloc. (Error code: 0)
- 批量预测UI报错:
Error: model server never became ready. Please validate that your model file or container configuration are valid.
我已尝试多种机器类型:
- c2-standard-16
- n1-highcpu-4
- n1-highcpu-16
但错误仍未改变。请问这是机器类型问题还是其他原因?该如何解决?
附批量预测的Cloud Function代码:
from google.cloud import aiplatform_v1beta1 from google.protobuf import json_format from google.protobuf.struct_pb2 import Value def hello_pubsub(event, context): project = "abc" display_name = "audience-segmentation-predictions" model_name = "projects/abc/locations/us-central1/models/19290" instances_format = "bigquery" bigquery_source_input_uri = "bq://abc.ml_final_g0_us.model_input_transformed_daily" predictions_format = "bigquery" bigquery_destination_output_uri = "bq://abc.ml_final_g0_us" location = "us-central1" api_endpoint = "us-central1-aiplatform.googleapis.com" def create_batch_prediction_job_bigquery_sample(project,display_name,model_name,instances_format,bigquery_source_input_uri,predictions_format,bigquery_destination_output_uri,location,api_endpoint): # The AI Platform services require regional API endpoints. client_options = {"api_endpoint": api_endpoint} # Initialize client that will be used to create and send requests. # This client only needs to be created once, and can be reused for multiple requests. client = aiplatform_v1beta1.JobServiceClient(client_options=client_options) model_parameters_dict = {} model_parameters = json_format.ParseDict(model_parameters_dict, Value()) batch_prediction_job = { "display_name": display_name, # Format: 'projects/{project}/locations/{location}/models/{model_id}' "model": model_name, "model_parameters": model_parameters, "input_config": { "instances_format": instances_format, "bigquery_source": {"input_uri": bigquery_source_input_uri}, }, "output_config": { "predictions_format": predictions_format, "bigquery_destination": {"output_uri": bigquery_destination_output_uri}, }, "instance_config": {"excluded_fields": "clientid"}, "dedicated_resources": { "machine_spec": { "machine_type": "c2-standard-30" }, }, } parent = f"projects/{project}/locations/{location}" response = client.create_batch_prediction_job( parent=parent, batch_prediction_job=batch_prediction_job ) print("response:", response) create_batch_prediction_job_bigquery_sample(project,display_name,model_name,instances_format,bigquery_source_input_uri,predictions_format,bigquery_destination_output_uri,location,api_endpoint)
std::bad_alloc错误本质是内存不足,虽然你尝试了不同CPU配置的机器,但可能忽略了内存资源的匹配,或者模型本身存在加载问题,以下是具体排查和解决步骤:
1. 优先调整内存资源配置
你之前选的机器类型内存配比偏低:
- n1-highcpu-4:仅3.6GB内存
- n1-highcpu-16:仅14.4GB内存
- c2-standard-16:64GB内存,但对较大的KMeans模型可能仍不够
建议:
- 改用高内存机型,比如
n1-highmem-16(16vCPU + 104GB内存)或c2-standard-60(60vCPU + 240GB内存),优先保证内存容量足够加载模型。 - 在
dedicated_resources中明确配置高内存机型,避免仅关注CPU核心数。
2. 验证模型文件完整性
- 本地尝试加载
model.joblib文件,确认模型本身无损坏:
如果本地加载也报错,说明模型文件在导出或上传时损坏,需要重新训练并导出模型。import joblib model = joblib.load("model.joblib") # 用少量测试数据运行预测,验证模型功能正常 - 查看
model.joblib的文件大小,如果超过几十GB,检查训练时是否存储了不必要的大对象(比如训练数据集副本),重新导出时只保留模型必要参数。
3. 拆分批量预测任务
将BigQuery中的30万条输入数据拆分为多个小批量任务运行,减少单任务的内存占用压力,比如按日期或用户ID分段处理。
4. 检查容器版本兼容性
Vertex AI运行scikit-learn模型使用官方容器,若训练时的scikit-learn版本与容器版本不一致,可能导致模型加载失败:
- 确认训练环境的scikit-learn版本,与Vertex AI批量预测容器的版本匹配。
- 若版本不兼容,可自定义容器镜像,使用和训练环境一致的依赖版本,再导入模型注册表进行批量预测。
内容的提问来源于stack exchange,提问作者Avantika Banerjee

