Spark2 Shell启动失败报java.lang.IllegalArgumentException: MALFORMED错误
java.lang.IllegalArgumentException: MALFORMED When Starting Spark2 Shell in CDH 5.14.2 Let's break down your problem and walk through practical steps to resolve it.
Context
You're running Cloudera CDH 5.14.2 with Java 1.8.0_91, and when trying to launch the Spark2 Shell, you hit a java.lang.IllegalArgumentException: MALFORMED error. You can't pinpoint exactly which JAR file is causing the issue. Here's your environment and error details:
Environment Version Output
Running spark2-shell --version gives:
$ spark2-shell --version Welcome to ____ __ / __/__ ___ _____/ /__ _\ \/ _ \/ _ `/ __/ '_/ /___/ .__/\_,_/_/ /_/\_\ version 2.2.0.cloudera1 /_/ Using Scala version 2.11.8, OpenJDK 64-Bit Server VM, 1.8.0_91 Branch HEAD Compiled by user jenkins on 2017-07-13T00:28:58Z Revision 39f5a2b89d29d5d420d88ce15c8c55e2b45aeb2e Url git://github.mtv.cloudera.com/CDH/spark.git Type --help for more information.
Startup Error Log
$ spark2-shell SLF4J: Class path contains multiple SLF4J bindings. SLF4J: Found binding in [jar:file:/usr/lib/zookeeper/lib/slf4j-log4j12-1.7.5.jar!/org/slf4j/impl/StaticLoggerBinder.class] SLF4J: Found binding in [jar:file:/usr/lib/flume-ng/lib/slf4j-log4j12-1.7.5.jar!/org/slf4j/impl/StaticLoggerBinder.class] SLF4J: Found binding in [jar:file:/usr/lib/parquet/lib/slf4j-log4j12-1.7.5.jar!/org/slf4j/impl/StaticLoggerBinder.class] SLF4J: See http://www.slf4j.org/codes.html#multiple_bindings for an explanation. SLF4J: Actual binding is of type [org.slf4j.impl.Log4jLoggerFactory] Exception in thread "main" java.lang.IllegalArgumentException: MALFORMED at java.util.zip.ZipCoder.toString(ZipCoder.java:58) at java.util.zip.ZipFile.getZipEntry(ZipFile.java:566) at java.util.zip.ZipFile.access$900(ZipFile.java:60) at java.util.zip.ZipFile$ZipEntryIterator.next(ZipFile.java:524) at java.util.zip.ZipFile$ZipEntryIterator.nextElement(ZipFile.java:499) at java.util.zip.ZipFile$ZipEntryIterator.nextElement(ZipFile.java:480) at scala.reflect.io.FileZipArchive.x$1$lzycompute(ZipArchive.scala:135) at scala.reflect.io.FileZipArchive.x$1(ZipArchive.scala:123) at scala.reflect.io.FileZipArchive.root$lzycompute(ZipArchive.scala:123) at scala.reflect.io.FileZipArchive.root(ZipArchive.scala:123) at scala.reflect.io.FileZipArchive.iterator(ZipArchive.scala:152) at scala.collection.IterableLike$class.foreach(IterableLike.scala:72) at scala.reflect.io.AbstractFile.foreach(AbstractFile.scala:91) at scala.tools.nsc.util.DirectoryClassPath.traverse(ClassPath.scala:277) at scala.tools.nsc.util.DirectoryClassPath.x$15$lzycompute(ClassPath.scala:299) at scala.tools.nsc.util.DirectoryClassPath.x$15(ClassPath.scala:299) at scala.tools.nsc.util.DirectoryClassPath.packages$lzycompute(ClassPath.scala:299) at scala.tools.nsc.util.DirectoryClassPath.packages(ClassPath.scala:299) at scala.tools.nsc.util.DirectoryClassPath.packages(ClassPath.scala:264) at scala.tools.nsc.util.MergedClassPath$$anonfun$packages$1.apply(ClassPath.scala:358) at scala.tools.nsc.util.MergedClassPath$$anonfun$packages$1.apply(ClassPath.scala:358) at scala.collection.Iterator$class.foreach(Iterator.scala:893) at scala.collection.AbstractIterator.foreach(Iterator.scala:1336) at scala.collection.IterableLike$class.foreach(IterableLike.scala:72) at scala.collection.AbstractIterable.foreach(Iterable.scala:54) at scala.tools.nsc.util.MergedClassPath.packages$lzycompute(ClassPath.scala:358) at scala.tools.nsc.util.MergedClassPath.packages(ClassPath.scala:353) at scala.tools.nsc.symtab.SymbolLoaders$PackageLoader$$anonfun$doComplete$1.apply$mcV$sp(SymbolLoaders.scala:269) at scala.tools.nsc.symtab.SymbolLoaders$PackageLoader$$anonfun$doComplete$1.apply(SymbolLoaders.scala:260) at scala.tools.nsc.symtab.SymbolLoaders$PackageLoader$$anonfun$doComplete$1.apply(SymbolLoaders.scala:260) at scala.reflect.internal.SymbolTable.enteringPhase(SymbolTable.scala:235) at scala.tools.nsc.symtab.SymbolLoaders$PackageLoader.doComplete(SymbolLoaders.scala:260) at scala.tools.nsc.symtab.SymbolLoaders$SymbolLoader.complete(SymbolLoaders.scala:211) at scala.reflect.internal.Symbols$Symbol.info(Symbols.scala:1514) at scala.reflect.internal.Mirrors$RootsBase.init(Mirrors.scala:256) at scala.tools.nsc.Global.rootMirror$lzycompute(Global.scala:73) at scala.tools.nsc.Global.rootMirror(Global.scala:71) at scala.tools.nsc.Global.rootMirror(Global.scala:39) at scala.reflect.internal.Definitions$DefinitionsClass.ObjectClass$lzycompute(Definitions.scala:257) at scala.reflect.internal.Definitions$DefinitionsClass.ObjectClass(Definitions.scala:257) at scala.reflect.internal.Definitions$DefinitionsClass.init(Definitions.scala:1394) at scala.tools.nsc.Global$Run.(Global.scala:1215) at scala.tools.nsc.interpreter.IMain.scala$tools$nsc$interpreter$IMain$$_initialize(IMain.scala:132) at scala.tools.nsc.interpreter.IMain.global$lzycompute(IMain.scala:161) at scala.tools.nsc.interpreter.IMain.global(IMain.scala:160) at scala.tools.nsc.interpreter.ILoop.command(ILoop.scala:680) at scala.tools.nsc.interpreter.ILoop.processLine(ILoop.scala:395) at org.apache.spark.repl.SparkILoop$$anonfun$initializeSpark$1.apply$mcV$sp(SparkILoop.scala:38) at org.apache.spark.repl.SparkILoop
What's Causing This?
The MALFORMED error comes from the JDK's ZipCoder class, which means Spark's Scala interpreter hit either a corrupted JAR file, or a JAR with entries/filenames using non-UTF-8 encoding, while scanning the classpath. Since the interpreter needs to traverse all JARs to build its symbol table, even one bad JAR will break the startup.
Step-by-Step Fixes
1. Locate the Corrupted JAR
Use this bash script to scan every JAR in Spark's classpath and identify the broken one:
# Get Spark2's full classpath SPARK_CLASSPATH=$(spark2-shell --verbose 2>&1 | grep "CLASSPATH" | cut -d'=' -f2) # Split the classpath and check each JAR IFS=':' read -ra JARS <<< "$SPARK_CLASSPATH" for jar in "${JARS[@]}"; do if [[ -f "$jar" ]]; then echo "Checking $jar..." unzip -t "$jar" > /dev/null 2>&1 if [[ $? -ne 0 ]]; then echo ">>> CORRUPTED JAR FOUND: $jar" fi fi done
Once you find the corrupted JAR, replace it with a working copy—either from another healthy node in your cluster or from the original CDH installation media.
2. Temporarily Exclude Suspicious JARs (Optional)
If you can't immediately find the bad JAR, try narrowing down the classpath. You can modify the spark2-shell startup script, or use the --jars flag to explicitly load only the JARs you need, avoiding the full default classpath.
3. Verify System Encoding
Java 8's ZipCoder is strict about encoding. Make sure your system uses UTF-8:
echo $LANG echo $LC_ALL
If not, set these variables temporarily and retry:
export LANG=en_US.UTF-8 export LC_ALL=en_US.UTF-8
4. Repair CDH-Managed JARs
If the corrupted JAR is part of the CDH distribution, use Cloudera Manager to redeploy the affected service. This will automatically replace any missing or damaged files with fresh copies from the CDH repository.
Final Notes
The root cause is almost always a corrupted JAR in the classpath. The scanning script will quickly point you to the culprit, and replacing it should resolve the issue. Double-checking the system encoding is also a quick win if the problem stems from non-UTF-8 filenames in JAR entries.
内容的提问来源于stack exchange,提问作者Rupesh More

