AWS EMR Step与EMR主节点spark-submit提交方式的差异、性能对比及最佳方案咨询
Great question! Let's break down the differences between these two Spark job submission approaches on AWS EMR, along with answers to your other follow-up questions:
Key Differences Between the Two Submission Methods
Let's walk through the core distinctions:
- Submission Context & Access
- The
aws emr add-stepscommand is a remote submission method: you don't need to SSH into the EMR master node at all. You can run this from your local machine, a CI/CD pipeline, or any server with AWS CLI credentials configured. - The
spark-submitcommand on the master node requires you to first SSH into the EMR cluster's master node. It's a direct, in-cluster submission.
- The
- Job Management & Visibility
- EMR Steps are managed by the EMR cluster's Step Manager. You can track their status directly in the AWS EMR console, set failure behaviors (like
ActionOnFailure=CONTINUEin your example), and get consolidated logs stored in S3 (if you configured cluster logging to S3). - Jobs submitted via
spark-submitare managed directly by YARN. To check status, you'll use YARN CLI commands likeyarn application -listoryarn logs -applicationId <id>. There's no EMR-level step tracking, so you don't get built-in failure actions at the cluster step level.
- EMR Steps are managed by the EMR cluster's Step Manager. You can track their status directly in the AWS EMR console, set failure behaviors (like
- Command Flexibility
- EMR Steps wrap Spark job arguments into the
Argsfield of the CLI command. You're limited to passing Spark-specific parameters here, and the structure is tied to EMR's step format. spark-submitgives you full access to all Spark submit options (like--conffor custom configurations,--executor-memory, etc.) directly in the command line, making it easier to tweak job parameters on the fly.
- EMR Steps wrap Spark job arguments into the
- Use Case Alignment
aws emr add-stepsis ideal for automated, scheduled jobs (e.g., triggered by Airflow, CloudWatch Events, or a CI/CD pipeline). It's designed for production-grade, unattended execution.spark-submiton the master node is perfect for development, debugging, and quick testing. You can modify code directly on the master node (or copy it via SCP) and submit it immediately without going through remote API calls.
Does spark-submit on the Master Node Run Distributedly?
Absolutely! When you run spark-submit --master yarn --deploy-mode cluster, the job will run across your EMR core nodes in a distributed way:
- The
--master yarnflag tells Spark to use YARN as the cluster manager, which allocates resources (executors) across your core nodes. - The
--deploy-mode clustermeans the Spark driver (the process that coordinates the job) runs on a node managed by YARN (this could be the master node or a core node, depending on YARN's scheduling logic). - All the actual data processing work happens on executors that are launched on your core nodes—your master node only acts as the submission entry point, not the execution engine.
Performance Comparison
In terms of job execution performance, the two methods are nearly identical because both ultimately submit the Spark job to YARN for execution. The only minor differences are:
- Submission Overhead:
aws emr add-stepshas a tiny bit of overhead from the AWS API call, but this is negligible for most jobs (it's just a few milliseconds of delay).spark-submiton the master node submits directly to YARN, so submission latency is slightly lower. - Log Handling: If your EMR cluster is configured to send Step logs to S3, there's a small amount of background IO to upload logs during/after the job, but this doesn't impact the actual processing performance of the Spark job itself.
- Resource Allocation: Both methods use YARN's resource scheduler, so the way executors are allocated and used is the same—no difference in compute efficiency.
Which is the Better Choice?
It depends on your use case:
- Go with
aws emr add-stepsif:- You're running production jobs that need to be automated or scheduled.
- You want centralized visibility and management via the EMR console.
- You need built-in failure handling (like terminating the cluster or continuing to the next step if a job fails).
- Go with
spark-submiton the master node if:- You're developing or debugging Spark code and need quick iteration.
- You want to test job configurations on the fly without setting up remote API calls.
- You're running ad-hoc, one-off jobs where you don't need long-term tracking or automation.
内容的提问来源于stack exchange,提问作者Deepak P R
相关产品推荐
相关产品推荐

