Vertex AI自定义训练CPU配额耗尽错误,但使用率未达上限
Vertex AI自定义任务配额超限问题排查
问题场景
使用以下Python脚本触发Vertex AI自定义任务:
print("Creating custom job...") job = aiplatform.CustomJob( display_name=f"process-text-{timestamp}", worker_pool_specs=[{ "machine_spec": { "machine_type": "n1-standard-4", }, "replica_count": 1, "container_spec": { "image_uri": TEXT_PROCESSOR_IMAGE, "args": [bucket_name, file_name] }, }] )
执行时触发以下错误:
grpc._channel._InactiveRpcError: <_InactiveRpcError of RPC that terminated with: status = StatusCode.RESOURCE_EXHAUSTED details = "The following quota metrics exceed quota limits: aiplatform.googleapis.com/custom_model_training_cpus" debug_error_string = "UNKNOWN:Error received from peer ipv4:142.250.152.95:443 {created_time:"2025-02-24T13:58:36.944174121+00:00", grpc_status:8, grpc_message:"The following quota metrics exceed quota limits: aiplatform.googleapis.com/custom_model_training_cpus"}" > The above exception was the direct cause of the following exception: raise exceptions.from_grpc_error(exc) from exc google.api_core.exceptions.ResourceExhausted: 429 The following quota metrics exceed quota limits: aiplatform.googleapis.com/custom_model_training_cpus
但在GCP控制台配额页面按「当前使用率百分比」排序后,所有资源使用率均未超过70%;申请额外配额环节也未发现接近限制的资源。
排查与解决方法
- 确认区域匹配:Vertex AI配额按区域划分,检查脚本中任务的部署区域(未指定时默认
us-central1)是否与控制台查看的区域一致。可在CustomJob初始化时显式指定location参数,例如location="us-central1",同时在控制台切换到对应区域查看配额。 - 清理残留任务:前往Vertex AI「自定义训练」页面,检查是否有运行中、排队中的未终止任务,手动清理这些任务释放CPU配额。
- 精准搜索配额指标:在控制台配额页面的搜索框直接输入
aiplatform.googleapis.com/custom_model_training_cpus,定位对应配额项查看实际使用与限制值,避免因指标名称差异导致漏看。 - 检查共享配额池:部分CPU配额为Vertex AI任务共享,确认是否有其他并行运行的Vertex AI任务占用了配额资源。
- 临时调整机器规格:紧急情况下,可降低任务的机器规格,例如将
n1-standard-4替换为n1-standard-2,减少单任务CPU占用以避开配额限制。
内容的提问来源于stack exchange,提问作者Evanss
相关产品推荐
相关产品推荐

