Spark与Kafka集成报错:Spark Submit出现NoClassDefFoundError求助
NoClassDefFoundError: org/apache/kafka/clients/consumer/Consumer in Spark Submit Hey there! I see you're hitting a common snag when moving from running Spark+Kafka code in Eclipse to using spark-submit—let's break this down and get it fixed.
Why this happens
When you run code in Eclipse, all your dependency jars (including Kafka clients) get loaded directly from your build setup. But spark-submit doesn’t automatically include every dependency by default, especially when there are version conflicts or incorrect scope settings in your build config.
Looking at your sbt configuration, I spot two key issues causing this error:
- Version conflict with
spark-streaming-kafka-assembly: You’ve included version 1.5.2 of this library, which is way older than your Spark 2.1.1 setup. This outdated jar clashes with your newer Spark and Kafka dependencies, leading to missing or incompatible classes. - Missing runtime dependencies: Even though you have
kafka-clientslisted, if you’re not packaging it into your jar or tellingspark-submitto fetch it, those classes won’t be available when your job runs on the cluster.
Step-by-step fixes
1. Clean up your sbt dependencies
First, remove the outdated spark-streaming-kafka-assembly line—it’s unnecessary with the newer spark-streaming-kafka-0-10 library. Here’s your corrected config:
name := "spark_streaming" version := "0.0.1" scalaVersion := "2.11.8" libraryDependencies ++= Seq( "org.apache.spark" %% "spark-core" % "2.1.1", "org.apache.spark" %% "spark-sql" % "2.1.1", "org.apache.spark" %% "spark-mllib" % "2.1.1", "org.apache.spark" %% "spark-hive" % "2.1.1", "org.apache.spark" %% "spark-streaming" % "2.1.1" % "provided", "org.apache.kafka" %% "kafka" % "0.11.0.0", "org.apache.spark" %% "spark-streaming-kafka-0-10" % "2.1.1", "org.apache.spark" %% "spark-sql-kafka-0-10" % "2.1.1", "org.apache.kafka" % "kafka-clients" % "0.11.0.0" )
2. Ensure dependencies are available at runtime
You have two straightforward options here:
Option 1: Build a "fat jar" with all dependencies
Use the sbt-assembly plugin to package your code along with all required dependencies (excluding theprovidedSpark libraries). Add this to yourproject/plugins.sbt:addSbtPlugin("com.eed3si9n" % "sbt-assembly" % "0.14.10")Run
sbt assemblyto generate the fat jar, then submit it withspark-submitas usual.Option 2: Let
spark-submitfetch dependencies
Skip building a fat jar and specify the required dependencies directly in your submit command. This handles dependency resolution for you and avoids bloating your jar:spark-submit \ --packages org.apache.spark:spark-streaming-kafka-0-10_2.11:2.1.1,org.apache.kafka:kafka-clients:0.11.0.0 \ --class your.package.MainClass \ path/to/your-regular-jar.jar
3. Double-check compatibility
Good news—your Spark 2.1.1 and Kafka 0.11.0.0 versions are compatible, and your Scala version (2.11.8) matches what Spark 2.1.1 uses. No mismatches here!
Final checks
After updating your config and choosing one of the dependency handling options, rebuild your project and re-run spark-submit. The NoClassDefFoundError should disappear because the Kafka client classes will now be available to your Spark application at runtime.
内容的提问来源于stack exchange,提问作者Guruprasad Swaminathan

