Google Cloud:如何在Cloud Datalab中用%%bash命令运行Cloud ML训练任务
Absolutely! You can use Datalab's %%bash magic command to submit training jobs directly to Cloud ML Engine (now part of Vertex AI's AI Platform). Here's how to make it work with your TensorFlow/Keras model:
First, Prepare Your Training Script
The code snippet you shared is a core part of your model, but Cloud ML needs a standalone, runnable Python script. Let's adapt it into something that can execute independently:
Save this as train_model.py (include all necessary imports and training logic):
import tensorflow as tf from tensorflow.contrib.layers import fully_connected def main(_): # Hyperparameters (you can pass these via command line too) num_inputs = 10 num_hidden = 32 num_outputs = 10 learning_rate = 0.001 epochs = 50 batch_size = 32 # Model definition (your existing code) X = tf.placeholder(tf.float32, shape=[None, num_inputs]) hidden = fully_connected(X, num_hidden, activation_fn=None) outputs = fully_connected(hidden, num_outputs, activation_fn=None) loss = tf.reduce_mean(tf.square(outputs - X)) optimizer = tf.train.AdamOptimizer(learning_rate) training_op = optimizer.minimize(loss) # Add training loop, data loading, and model saving to GCS with tf.Session() as sess: sess.run(tf.global_variables_initializer()) # Replace with your actual data pipeline for epoch in range(epochs): # Training steps here... print(f"Epoch {epoch+1} completed") # Save model to Cloud Storage (critical for Cloud ML) saver = tf.train.Saver() saver.save(sess, "gs://your-bucket-name/models/autoencoder") if __name__ == "__main__": tf.app.run()
Copy Your Script to Cloud Storage
Cloud ML needs access to your code, so copy the script to a GCS bucket using %%bash:
%%bash gsutil cp train_model.py gs://your-bucket-name/training-code/
Submit the Training Job via %%bash
Use the gcloud ai-platform jobs submit training command (adjust parameters to match your setup):
%%bash gcloud ai-platform jobs submit training my_cloud_ml_job_$(date +%Y%m%d_%H%M%S) \ --region us-central1 \ --runtime-version 1.15 \ --python-version 3.7 \ --job-dir gs://your-bucket-name/job-output/ \ --package-path gs://your-bucket-name/training-code/ \ --module-name train_model
Key Notes to Remember
- Permissions: Ensure your Datalab instance's service account has the
AI Platform Job Editorrole (or equivalent) to submit jobs. - Runtime Versions: Match the TensorFlow/Keras version in your script to the
--runtime-versionspecified (check Cloud ML's supported versions). - GCS Paths: All output (models, logs) should go to Cloud Storage—local paths won't persist on Cloud ML's training instances.
- Hyperparameters: You can pass hyperparameters to your script using the
--flag at the end of the gcloud command, then parse them in your Python code withargparse.
If you run into issues, check the job logs in the Cloud Console's AI Platform section—they'll help debug any runtime errors or permission issues.
内容的提问来源于stack exchange,提问作者Nicky Feller

