You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark初始作业无法获取资源:Amazon EC2独立集群问题

Hey Joe, let's troubleshoot why your Spark jobs can't acquire resources on your EC2 standalone cluster. I've dealt with similar setup hiccups before, so here are the key checks and fixes to work through:

1. Confirm Worker Nodes Are Registered with the Master

First, rule out if the Workers are even connecting to the Master:

  • Check the Master's Web UI (default port 8080) by visiting http://<master-public-dns>:8080. If no Worker nodes show up in the "Workers" section, they aren't registering successfully.
  • Jump into the Worker nodes' Spark logs (usually in ${SPARK_HOME}/logs/spark-*-worker-*.out). Look for errors like failed connection attempts to the Master or SSH authentication issues.
  • Test network connectivity from Workers to the Master: Run telnet <master-public-dns> 7077 or nc -zv <master-public-dns> 7077 on each Worker. If this fails, your EC2 security groups or network routing are blocking the Spark cluster communication port.

2. Validate the slaves File and SSH Setup

Your slaves file and SSH config are critical for starting Workers:

  • Double-check the slaves file: Ensure each Worker's public IP is on its own line, no extra spaces or typos. Also confirm every node uses the exact same SPARK_HOME path—the start-slaves.sh script relies on this to run commands remotely.
  • Verify passwordless SSH works: From the Master, run ssh <worker-public-ip> for each Worker. If you get a password prompt or permission error, fix the ~/.ssh/authorized_keys file on Workers (it must have 600 permissions) and ensure the SSH service is running on all nodes.
  • Run start-slaves.sh manually and watch the output: If you see "Connection refused" or "Permission denied" messages, that's a direct clue about SSH or network issues.

3. Tweak Spark Master/Worker Configuration

Incorrect network binding is a common culprit:

  • When starting the Master, did you explicitly set the public DNS with the --host flag? Like:
    ${SPARK_HOME}/sbin/start-master.sh --host <master-public-dns>
    
    If not, the Master might bind to the EC2 instance's private IP, which Workers (connecting via public IP) can't reach.
  • On Worker nodes, add SPARK_MASTER_HOST to spark-env.sh:
    export SPARK_MASTER_HOST=<master-public-dns>
    
    Without this, Workers might default to trying to connect to localhost, which obviously won't work.
  • Audit EC2 security groups:
    • Master node: Open ports 7077 (cluster communication), 8080 (UI), and 22 (SSH) to all Worker public IPs.
    • Worker nodes: Open port 8081 (UI) and 22 (SSH) to the Master's public IP.

4. Check Job Submission Parameters

Make sure your job is actually targeting the cluster:

  • When submitting, explicitly specify the Master address:
    spark-submit --master spark://<master-public-dns>:7077 --class your.main.Class your-app.jar
    
    If you omit this or use local[*], the job will run locally and ignore the cluster.
  • Verify resource requests match Worker capacity: If you request --total-executor-cores 8 but each Worker only has 2 cores, the job will hang waiting for unavailable resources. Check the Master UI to see each Worker's available cores/memory.

5. Dive Into Detailed Logs for Clues

If all else fails, dig deeper into logs:

  • Master logs (${SPARK_HOME}/logs/spark-*-master-*.out): Look for entries about job submissions and resource allocation failures—you might see messages like "No available executors to run task".
  • Driver logs: If submitting in client mode, logs will be in your terminal. In cluster mode, check the Master or Worker logs for the driver's output. Look for exceptions related to resource negotiation.

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:46:00