You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google ML Engine REST API提交训练任务遇错求助

Troubleshooting Google Cloud ML Engine Training Failures for Census Sample via REST API

Let's break down and solve your two-stage training failure problem step by step:

1. Fixing the "Required Arguments Missing" Error (Exit Code 2)

Your initial REST API args format was the root cause here. When passing command-line arguments via ML Engine's REST API, each flag and its corresponding value must be separate items in the args list—you can't combine them into a single string like you would in a shell command.

Wrong format (what you started with):

["--train_files gs://MY_BUCKET/adult.data.csv", "--eval_files gs://MY_BUCKET/adult.test.csv"]

Correct format (the fix you implemented):

["--train-files", "gs://MY_BUCKET/adult.data.csv", "--eval-files", "gs://MY_BUCKET/adult.test.csv", "--train-steps", "100", "--eval-steps", "10"]

This matches how the CLI parses arguments (splitting on spaces), so ML Engine can correctly identify each required flag and its value. This resolved the task.py: error: the following arguments are required and exit code 2 errors.

2. Debugging the "Exit Code 1" Error (Partial Training Completion)

Exit code 1 typically points to an application-level error (not a parameter or infrastructure issue) that occurs after training starts. Since you mentioned training ran partially, here are the most likely fixes to investigate:

  • Verify GCS Bucket Permissions
    ML Engine uses a service account (usually [PROJECT_NUMBER]-compute@developer.gserviceaccount.com) to access your GCS resources. Ensure this account has the following IAM roles on your bucket:

    • Storage Object Creator (to write model checkpoints, logs, and outputs)
    • Storage Object Viewer (to read training/evaluation data)
      Double-check that you haven't restricted bucket access to specific users or IPs that exclude the ML Engine service account.
  • Check Job Directory Configuration
    If your Census sample code uses the --job-dir flag (common for saving models/outputs), confirm:

    • The path is a valid GCS bucket/prefix (e.g., gs://MY_BUCKET/census-training-job)
    • The bucket exists and the service account can write to it
      Omitting --job-dir might cause the code to default to a local path that isn't persisted, leading to failures when it tries to save data.
  • Dig into Detailed Logs
    The short error message only tells you the master replica failed—you need to look at the full logs in Cloud Logging to find the root cause:

    1. Go to the Cloud Logging console and filter for your ML Engine job (use the job name as a filter)
    2. Search for keywords like PermissionDenied, FileNotFoundError, Exception, or Failed to write
    3. Look for stack traces that show exactly where the code crashed (e.g., during model saving or evaluation data loading)
  • Validate Evaluation Data and Code Logic
    Even if training ran, the failure could happen during evaluation:

    • Confirm adult.test.csv has the same format as adult.data.csv (no missing columns, correct delimiters)
    • Check the Census sample's model_fn or evaluation logic for edge cases (e.g., handling missing values, categorical features that aren't present in the test set but exist in training)

内容的提问来源于stack exchange,提问作者Johnathan Brown

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:38:38