You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS EMR Step与EMR主节点spark-submit提交方式的差异、性能对比及最佳方案咨询

Great question! Let's break down the differences between these two Spark job submission approaches on AWS EMR, along with answers to your other follow-up questions:

Key Differences Between the Two Submission Methods

Let's walk through the core distinctions:

  • Submission Context & Access
    • The aws emr add-steps command is a remote submission method: you don't need to SSH into the EMR master node at all. You can run this from your local machine, a CI/CD pipeline, or any server with AWS CLI credentials configured.
    • The spark-submit command on the master node requires you to first SSH into the EMR cluster's master node. It's a direct, in-cluster submission.
  • Job Management & Visibility
    • EMR Steps are managed by the EMR cluster's Step Manager. You can track their status directly in the AWS EMR console, set failure behaviors (like ActionOnFailure=CONTINUE in your example), and get consolidated logs stored in S3 (if you configured cluster logging to S3).
    • Jobs submitted via spark-submit are managed directly by YARN. To check status, you'll use YARN CLI commands like yarn application -list or yarn logs -applicationId <id>. There's no EMR-level step tracking, so you don't get built-in failure actions at the cluster step level.
  • Command Flexibility
    • EMR Steps wrap Spark job arguments into the Args field of the CLI command. You're limited to passing Spark-specific parameters here, and the structure is tied to EMR's step format.
    • spark-submit gives you full access to all Spark submit options (like --conf for custom configurations, --executor-memory, etc.) directly in the command line, making it easier to tweak job parameters on the fly.
  • Use Case Alignment
    • aws emr add-steps is ideal for automated, scheduled jobs (e.g., triggered by Airflow, CloudWatch Events, or a CI/CD pipeline). It's designed for production-grade, unattended execution.
    • spark-submit on the master node is perfect for development, debugging, and quick testing. You can modify code directly on the master node (or copy it via SCP) and submit it immediately without going through remote API calls.
Does spark-submit on the Master Node Run Distributedly?

Absolutely! When you run spark-submit --master yarn --deploy-mode cluster, the job will run across your EMR core nodes in a distributed way:

  • The --master yarn flag tells Spark to use YARN as the cluster manager, which allocates resources (executors) across your core nodes.
  • The --deploy-mode cluster means the Spark driver (the process that coordinates the job) runs on a node managed by YARN (this could be the master node or a core node, depending on YARN's scheduling logic).
  • All the actual data processing work happens on executors that are launched on your core nodes—your master node only acts as the submission entry point, not the execution engine.
Performance Comparison

In terms of job execution performance, the two methods are nearly identical because both ultimately submit the Spark job to YARN for execution. The only minor differences are:

  • Submission Overhead: aws emr add-steps has a tiny bit of overhead from the AWS API call, but this is negligible for most jobs (it's just a few milliseconds of delay). spark-submit on the master node submits directly to YARN, so submission latency is slightly lower.
  • Log Handling: If your EMR cluster is configured to send Step logs to S3, there's a small amount of background IO to upload logs during/after the job, but this doesn't impact the actual processing performance of the Spark job itself.
  • Resource Allocation: Both methods use YARN's resource scheduler, so the way executors are allocated and used is the same—no difference in compute efficiency.
Which is the Better Choice?

It depends on your use case:

  • Go with aws emr add-steps if:
    • You're running production jobs that need to be automated or scheduled.
    • You want centralized visibility and management via the EMR console.
    • You need built-in failure handling (like terminating the cluster or continuing to the next step if a job fails).
  • Go with spark-submit on the master node if:
    • You're developing or debugging Spark code and need quick iteration.
    • You want to test job configurations on the fly without setting up remote API calls.
    • You're running ad-hoc, one-off jobs where you don't need long-term tracking or automation.

内容的提问来源于stack exchange,提问作者Deepak P R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 00:54:05