如何通过Terraform创建EMR时添加步骤并禁用keep_job_flow_alive_when_no_steps
Got it, let's walk through exactly how to set this up with Terraform. When you disable keep_job_flow_alive_when_no_steps, your EMR cluster will shut down right after all its steps finish running. The critical thing here is making sure your steps are tied to the cluster correctly so it doesn't terminate immediately before the steps even start.
keep_job_flow_alive_when_no_steps = false First, define your EMR cluster resource and explicitly set the flag to disable keeping the cluster alive without steps. This is the core setting that triggers termination once steps complete.
You have two main ways to attach steps to your cluster—choose the one that fits your workflow:
Approach 1: Embed Steps Directly in the Cluster Resource
This is the most straightforward and reliable method, as steps are registered at the same time the cluster is created. The cluster will start executing steps immediately after launching, and terminate once all steps finish.
Here's a complete example:
resource "aws_emr_cluster" "my_emr_cluster" { name = "my-emr-cluster-with-embedded-steps" release_label = "emr-6.15.0" # Use a version compatible with your applications applications = ["Spark"] # Add other apps like Hive if needed instances { instance_group { instance_type = "m5.xlarge" instance_count = 1 instance_role = "MASTER" } instance_group { instance_type = "m5.xlarge" instance_count = 2 instance_role = "CORE" } # Disable keeping the cluster alive when no steps are running keep_job_flow_alive_when_no_steps = false } # Embed your steps directly in the cluster configuration steps { name = "Spark Pi Calculation" action_on_failure = "TERMINATE_CLUSTER" # Adjust based on your failure needs hadoop_jar_step { jar = "command-runner.jar" # EMR's built-in command runner args = [ "spark-submit", "--class", "org.apache.spark.examples.SparkPi", "--master", "yarn", "--deploy-mode", "cluster", "local:///usr/lib/spark/examples/jars/spark-examples.jar", "10" # Number of iterations for the Pi calculation ] } } # Replace these with your existing IAM roles or create them in Terraform service_role = aws_iam_role.emr_service_role.arn job_flow_role = aws_iam_role.emr_instance_role.arn }
Approach 2: Use Separate aws_emr_step Resources
If you need to manage steps independently (e.g., dynamic step generation, separate lifecycle management), you can create standalone step resources. However, you need to be careful with timing—if the cluster terminates before the step is submitted, it will fail.
Here's how to do it safely:
resource "aws_emr_cluster" "my_emr_cluster" { name = "my-emr-cluster-separate-steps" release_label = "emr-6.15.0" applications = ["Spark"] instances { instance_group { instance_type = "m5.xlarge" instance_count = 1 instance_role = "MASTER" } instance_group { instance_type = "m5.xlarge" instance_count = 2 instance_role = "CORE" } keep_job_flow_alive_when_no_steps = false } service_role = aws_iam_role.emr_service_role.arn job_flow_role = aws_iam_role.emr_instance_role.arn } # Define your step as a separate resource with explicit dependency resource "aws_emr_step" "spark_pi_step" { cluster_id = aws_emr_cluster.my_emr_cluster.id name = "Spark Pi Calculation" action_on_failure = "TERMINATE_CLUSTER" hadoop_jar_step { jar = "command-runner.jar" args = [ "spark-submit", "--class", "org.apache.spark.examples.SparkPi", "--master", "yarn", "--deploy-mode", "cluster", "local:///usr/lib/spark/examples/jars/spark-examples.jar", "10" ] } # Ensure the step is created only after the cluster is ready depends_on = [aws_emr_cluster.my_emr_cluster] }
Critical Note for Approach 2:
There's a small window where the cluster might terminate before the step is submitted. To avoid this, you can:
- Initially set
keep_job_flow_alive_when_no_steps = true, apply to create the cluster and steps. - Update the flag to
falseand re-apply—this will trigger the cluster to terminate once steps finish.
- Action on Failure: Adjust the
action_on_failureparameter based on your needs:TERMINATE_CLUSTER: Shuts down the cluster immediately if the step fails.CONTINUE: Moves to the next step even if this one fails.CANCEL_AND_WAIT: Stops subsequent steps but keeps the cluster running (less useful withkeep_job_flow_alive_when_no_steps = false).
- IAM Permissions: Make sure your EMR service role and instance profile have permissions to access any resources your steps need (e.g., S3 buckets for input/output data).
- Cluster Termination: Once all steps complete (or fail with
TERMINATE_CLUSTER), the cluster will automatically terminate—no manual cleanup needed.
内容的提问来源于stack exchange,提问作者Nipun

