Spark启动异常求助:执行pyspark --master yarn --num-executors 3时报错Java gateway process exited before sending its port number
Looks like your local PySpark setup works perfectly fine, but you're hitting a roadblock when switching to YARN mode. That vague "Java gateway process exited" error usually points to communication issues between the PySpark frontend and the Java backend, or missing configurations specific to YARN. Let's walk through actionable steps to fix this:
1. Double-check your Java environment variables
Even though you have Java 8 installed, Spark might not be picking it up correctly:
- Run
echo $JAVA_HOMEin your terminal. It should point directly to your Java 8 installation (e.g.,/Library/Java/JavaVirtualMachines/jdk1.8.0_301.jdk/Contents/Homeon macOS). - If
JAVA_HOMEisn't set, add these lines to your shell config file (.bash_profile,.zshrc, etc.):export JAVA_HOME=/path/to/your/jdk1.8.0_XXX export PATH=$JAVA_HOME/bin:$PATH - Restart your terminal and confirm with
java -versionthat you're indeed running Java 8.
2. Ensure YARN is running and reachable
Spark can't connect to a YARN cluster that's offline:
- If you're using a local single-node YARN setup, start it first with
start-yarn.sh. - For a remote cluster, verify YARN services are up: run
yarn node -listvia the CLI, or check the YARN ResourceManager UI (typically athttp://<resource-manager-host>:8088).
3. Configure Spark to recognize YARN
Spark needs Hadoop/YARN config files to connect properly:
- Head to your Spark
confdirectory (usually$SPARK_HOME/conf). - Copy the template configs if you haven't already:
cp $SPARK_HOME/conf/spark-env.sh.template $SPARK_HOME/conf/spark-env.sh cp $SPARK_HOME/conf/spark-defaults.conf.template $SPARK_HOME/conf/spark-defaults.conf - In
spark-env.sh, add this line to point Spark to your Hadoop configs:export HADOOP_CONF_DIR=/path/to/your/hadoop/conf - In
spark-defaults.conf, add these basic YARN-related settings:spark.master yarn spark.driver.memory 1g spark.executor.memory 1g
4. Get detailed logs to find the root cause
The default error message doesn't tell the whole story—enable debug logging to see what's failing:
- Run PySpark with verbose logging:
pyspark --master yarn --num-executors 3 --conf spark.driver.logLevel=DEBUG - Check the terminal output for exceptions like missing dependencies, permission errors, or memory allocation failures.
- If you can get an application ID from the error, pull YARN logs with:
These logs often reveal exactly why the Java gateway crashed.yarn logs -applicationId <your-app-id>
5. Verify file permissions
Spark might fail to access critical files needed to start the Java gateway:
- Ensure your user has read access to
$HADOOP_CONF_DIRand all files inside it. - On shared clusters, confirm you have permission to submit YARN applications (check with your cluster admin if unsure).
6. Check Spark-Hadoop compatibility
Spark 3.1.2 works best with Hadoop 2.7 and above. Make sure your YARN/Hadoop version is compatible with your Spark release—mismatches can cause silent failures in the Java gateway.
Go through these steps one by one, and try running the YARN command again. If you spot specific errors in the logs, that'll help narrow down the fix even faster!
内容的提问来源于stack exchange,提问作者Ingo Banghard

