Spark 2.1.0 YARN集群运行SparkPi时遇NoSuchMethodError问题求助
NoSuchMethodError with Spark 2.1.0 on YARN Cluster Mode (Hadoop 2.7.4) Hey there! Let's break down and fix this issue you're hitting. That java.lang.NoSuchMethodError for org.apache.hadoop.yarn.util.Apps.crossPlatformify boils down to a version conflict between Spark's bundled Hadoop dependencies and your cluster's Hadoop installation—and it only pops up in YARN cluster mode for a specific reason.
Why This Happens
- Spark 2.1.0 pre-built for Hadoop 2.7 uses Hadoop 2.7.3 under the hood, which includes the missing
crossPlatformifymethod. - Your cluster runs Hadoop 2.7.4, which doesn't include this method (or the method signature was altered).
- In client mode, the Spark Driver runs on your local machine, where Spark's bundled jars take priority in the classpath—so no conflict arises.
- In cluster mode, the Driver runs on a YARN node, and YARN's default classpath prioritizes the cluster's Hadoop jars over Spark's bundled ones. This means the runtime tries to use Hadoop 2.7.4's
hadoop-common.jar, which lacks the required method.
Solutions to Try
1. Force Spark's Bundled Jars to Take Priority
Configure Spark to use its own bundled jars instead of the cluster's Hadoop jars when running on YARN. You can do this either per-submit or via a global config:
Per-Submit Command
Add the spark.yarn.jars configuration to your spark-submit command (the path must exist on all cluster nodes):
${SPARK_HOME}/bin/spark-submit \ --class org.apache.spark.examples.SparkPi \ --master yarn \ --deploy-mode cluster \ --driver-memory 1g \ --conf spark.yarn.jars=file:///opt/spark-2.1.0-bin-hadoop2.7/jars/*.jar \ lib/spark-examples*.jar 10
Global Config (spark-defaults.conf)
Edit ${SPARK_HOME}/conf/spark-defaults.conf to add this line, so all future submissions use this setting automatically:
spark.yarn.jars file:///opt/spark-2.1.0-bin-hadoop2.7/jars/*.jar
2. Adjust YARN's Application Classpath
Modify the YARN classpath to prioritize Spark's jars over the cluster's Hadoop jars. Edit yarn-site.xml (usually at /opt/hadoop-2.7.4/etc/hadoop/yarn-site.xml) on all nodes:
Find the yarn.application.classpath property, and prepend Spark's jars directory to the list:
<property> <name>yarn.application.classpath</name> <value> /opt/spark-2.1.0-bin-hadoop2.7/jars/*, /opt/hadoop-2.7.4/share/hadoop/common/*, /opt/hadoop-2.7.4/share/hadoop/common/lib/*, <!-- Keep the rest of your existing classpath entries here --> </value> </property>
Restart YARN services on all nodes after making this change for it to take effect.
3. Explicitly Set Driver/Executor Classpaths
If the above fixes don't work, explicitly point the driver and executor classpaths to the correct hadoop-common-2.7.3.jar:
${SPARK_HOME}/bin/spark-submit \ --class org.apache.spark.examples.SparkPi \ --master yarn \ --deploy-mode cluster \ --driver-memory 1g \ --conf spark.driver.extraClassPath=/opt/spark-2.1.0-bin-hadoop2.7/jars/hadoop-common-2.7.3.jar \ --conf spark.executor.extraClassPath=/opt/spark-2.1.0-bin-hadoop2.7/jars/hadoop-common-2.7.3.jar \ lib/spark-examples*.jar 10
Since the jar exists on all nodes, this will ensure the runtime uses the version with the missing method.
4. Recompile Spark for Hadoop 2.7.4 (Long-Term Fix)
For full, permanent compatibility, recompile Spark 2.1.0 against your exact Hadoop 2.7.4 version:
- Download the Spark 2.1.0 source code.
- Run the compilation command:
./dev/make-distribution.sh --name hadoop2.7.4 --tgz -Phadoop-2.7 -Dhadoop.version=2.7.4 - Replace your existing Spark installation with the newly compiled version.
Why --jars and --driver-library-path Didn't Work
--jarsadds jars to the executor classpath, but in cluster mode, the Driver's classpath is controlled by YARN's default settings—so this doesn't fix the Driver-side conflict.--driver-library-pathis for native libraries, not Java classpath entries, so it can't resolve this method-missing error.
内容的提问来源于stack exchange,提问作者user9158993

