Spark集群作业未分发至工作节点,仅在主节点运行求助
Hey there, let's figure out why your Spark job is only running on the master node (h1) instead of distributing to h2 and h3. I’ve tackled this exact issue plenty of times, so let’s walk through the most likely fixes step by step:
If your worker nodes aren’t registered with the master, Spark can’t send tasks to them. Here’s how to check:
- Fire up the Spark Master UI on h1 (default URL:
http://h1:8080). Look for the "Workers" section—do h2 and h3 show up here?- If they’re missing, head to h2 and h3’s Spark log directory (usually
$SPARK_HOME/logs) and check the worker logs (filenames likespark-ubuntu-org.apache.spark.deploy.worker.Worker-1-h2.out). You’ll often find errors here: network connectivity issues, port conflicts, or mismatched Spark configurations between master and workers. - Double-check how you started the workers: you need to point them directly to the master with a command like
./sbin/start-worker.sh spark://h1:7077(7077 is Spark’s default master port). If you used the wrong master address, workers won’t register.
- If they’re missing, head to h2 and h3’s Spark log directory (usually
Your current submit command doesn’t specify a master URL, which means Spark might be falling back to local mode (only running on the machine you submit from).
- Update your command to explicitly target the Spark cluster master:
bin/spark-submit --master spark://h1:7077 --class org.dataalgorithms.chap07.spark.FindAssociationRules /home/ubuntu/project_spark/data-algorithms-1.0.0.jar ./in/xaa - Also, check your
$SPARK_HOME/conf/spark-defaults.conffile. If there’s a line likespark.master local[*]orspark.master local, that forces local mode by default—either remove that line or set it tospark.master spark://h1:7077.
Sometimes workers don’t get tasks because they don’t have enough resources to run executors.
- Check the "Executors" tab in the Spark UI (once you submit the job, the UI URL will show up in your terminal). Do you see executors from h2 and h3 here?
- If not, check master or worker logs for errors about insufficient memory/CPU. You might be requesting more executor memory than your workers have available. Try reducing the executor memory in your submit command:
bin/spark-submit --master spark://h1:7077 --executor-memory 1g --class org.dataalgorithms.chap07.spark.FindAssociationRules /home/ubuntu/project_spark/data-algorithms-1.0.0.jar ./in/xaa - Double-check
spark-defaults.conffor settings likespark.executor.memoryorspark.cores.max—make sure they don’t exceed the available resources on h2 and h3.
- If not, check master or worker logs for errors about insufficient memory/CPU. You might be requesting more executor memory than your workers have available. Try reducing the executor memory in your submit command:
If your input file ./in/xaa only exists on h1’s local filesystem, Spark will keep all tasks on h1 (it’s more efficient to run tasks where the data lives, instead of copying data across nodes).
- The fix here is to use HDFS (since you have a Hadoop namenode set up). Upload your data to HDFS first:
Then update your job’s input path to the HDFS URL:hdfs dfs -put ./in/xaa /in/xaabin/spark-submit --master spark://h1:7077 --class org.dataalgorithms.chap07.spark.FindAssociationRules /home/ubuntu/project_spark/data-algorithms-1.0.0.jar hdfs://h1:9000/in/xaa - If you must use local files, ensure
./in/xaaexists in the exact same path on h2 and h3—otherwise workers can’t access the data and Spark won’t send tasks there.
It’s rare, but sometimes hardcoded settings in your application override cluster configurations. Check the FindAssociationRules class’s SparkContext initialization:
- Bad (forces local mode):
SparkConf conf = new SparkConf().setAppName("FindAssociationRules").setMaster("local"); JavaSparkContext sc = new JavaSparkContext(conf); - Good (lets the submit command set the master):
SparkConf conf = new SparkConf().setAppName("FindAssociationRules"); JavaSparkContext sc = new JavaSparkContext(conf);
Start with checking the cluster connection and submit command—those are the most common culprits. Let me know if any of these steps get your jobs running on h2 and h3!
内容的提问来源于stack exchange,提问作者learning_dev

