Spark初始作业无法获取资源:Amazon EC2独立集群问题
Hey Joe, let's troubleshoot why your Spark jobs can't acquire resources on your EC2 standalone cluster. I've dealt with similar setup hiccups before, so here are the key checks and fixes to work through:
1. Confirm Worker Nodes Are Registered with the Master
First, rule out if the Workers are even connecting to the Master:
- Check the Master's Web UI (default port 8080) by visiting
http://<master-public-dns>:8080. If no Worker nodes show up in the "Workers" section, they aren't registering successfully. - Jump into the Worker nodes' Spark logs (usually in
${SPARK_HOME}/logs/spark-*-worker-*.out). Look for errors like failed connection attempts to the Master or SSH authentication issues. - Test network connectivity from Workers to the Master: Run
telnet <master-public-dns> 7077ornc -zv <master-public-dns> 7077on each Worker. If this fails, your EC2 security groups or network routing are blocking the Spark cluster communication port.
2. Validate the slaves File and SSH Setup
Your slaves file and SSH config are critical for starting Workers:
- Double-check the
slavesfile: Ensure each Worker's public IP is on its own line, no extra spaces or typos. Also confirm every node uses the exact sameSPARK_HOMEpath—thestart-slaves.shscript relies on this to run commands remotely. - Verify passwordless SSH works: From the Master, run
ssh <worker-public-ip>for each Worker. If you get a password prompt or permission error, fix the~/.ssh/authorized_keysfile on Workers (it must have600permissions) and ensure the SSH service is running on all nodes. - Run
start-slaves.shmanually and watch the output: If you see "Connection refused" or "Permission denied" messages, that's a direct clue about SSH or network issues.
3. Tweak Spark Master/Worker Configuration
Incorrect network binding is a common culprit:
- When starting the Master, did you explicitly set the public DNS with the
--hostflag? Like:
If not, the Master might bind to the EC2 instance's private IP, which Workers (connecting via public IP) can't reach.${SPARK_HOME}/sbin/start-master.sh --host <master-public-dns> - On Worker nodes, add
SPARK_MASTER_HOSTtospark-env.sh:
Without this, Workers might default to trying to connect toexport SPARK_MASTER_HOST=<master-public-dns>localhost, which obviously won't work. - Audit EC2 security groups:
- Master node: Open ports 7077 (cluster communication), 8080 (UI), and 22 (SSH) to all Worker public IPs.
- Worker nodes: Open port 8081 (UI) and 22 (SSH) to the Master's public IP.
4. Check Job Submission Parameters
Make sure your job is actually targeting the cluster:
- When submitting, explicitly specify the Master address:
If you omit this or usespark-submit --master spark://<master-public-dns>:7077 --class your.main.Class your-app.jarlocal[*], the job will run locally and ignore the cluster. - Verify resource requests match Worker capacity: If you request
--total-executor-cores 8but each Worker only has 2 cores, the job will hang waiting for unavailable resources. Check the Master UI to see each Worker's available cores/memory.
5. Dive Into Detailed Logs for Clues
If all else fails, dig deeper into logs:
- Master logs (
${SPARK_HOME}/logs/spark-*-master-*.out): Look for entries about job submissions and resource allocation failures—you might see messages like "No available executors to run task". - Driver logs: If submitting in client mode, logs will be in your terminal. In cluster mode, check the Master or Worker logs for the driver's output. Look for exceptions related to resource negotiation.
内容的提问来源于stack exchange,提问作者Joe
相关产品推荐
相关产品推荐

