Google ML Engine REST API提交训练任务遇错求助
Let's break down and solve your two-stage training failure problem step by step:
1. Fixing the "Required Arguments Missing" Error (Exit Code 2)
Your initial REST API args format was the root cause here. When passing command-line arguments via ML Engine's REST API, each flag and its corresponding value must be separate items in the args list—you can't combine them into a single string like you would in a shell command.
Wrong format (what you started with):
["--train_files gs://MY_BUCKET/adult.data.csv", "--eval_files gs://MY_BUCKET/adult.test.csv"]
Correct format (the fix you implemented):
["--train-files", "gs://MY_BUCKET/adult.data.csv", "--eval-files", "gs://MY_BUCKET/adult.test.csv", "--train-steps", "100", "--eval-steps", "10"]
This matches how the CLI parses arguments (splitting on spaces), so ML Engine can correctly identify each required flag and its value. This resolved the task.py: error: the following arguments are required and exit code 2 errors.
2. Debugging the "Exit Code 1" Error (Partial Training Completion)
Exit code 1 typically points to an application-level error (not a parameter or infrastructure issue) that occurs after training starts. Since you mentioned training ran partially, here are the most likely fixes to investigate:
Verify GCS Bucket Permissions
ML Engine uses a service account (usually[PROJECT_NUMBER]-compute@developer.gserviceaccount.com) to access your GCS resources. Ensure this account has the following IAM roles on your bucket:Storage Object Creator(to write model checkpoints, logs, and outputs)Storage Object Viewer(to read training/evaluation data)
Double-check that you haven't restricted bucket access to specific users or IPs that exclude the ML Engine service account.
Check Job Directory Configuration
If your Census sample code uses the--job-dirflag (common for saving models/outputs), confirm:- The path is a valid GCS bucket/prefix (e.g.,
gs://MY_BUCKET/census-training-job) - The bucket exists and the service account can write to it
Omitting--job-dirmight cause the code to default to a local path that isn't persisted, leading to failures when it tries to save data.
- The path is a valid GCS bucket/prefix (e.g.,
Dig into Detailed Logs
The short error message only tells you the master replica failed—you need to look at the full logs in Cloud Logging to find the root cause:- Go to the Cloud Logging console and filter for your ML Engine job (use the job name as a filter)
- Search for keywords like
PermissionDenied,FileNotFoundError,Exception, orFailed to write - Look for stack traces that show exactly where the code crashed (e.g., during model saving or evaluation data loading)
Validate Evaluation Data and Code Logic
Even if training ran, the failure could happen during evaluation:- Confirm
adult.test.csvhas the same format asadult.data.csv(no missing columns, correct delimiters) - Check the Census sample's
model_fnor evaluation logic for edge cases (e.g., handling missing values, categorical features that aren't present in the test set but exist in training)
- Confirm
内容的提问来源于stack exchange,提问作者Johnathan Brown

