Scala2.11+sbt1.0环境下Spark Streaming Kafka报ClassNotFoundException求助
java.lang.ClassNotFoundException for Spark Streaming + Kafka Setup Hey there, let's work through that java.lang.ClassNotFoundException you're facing with your Scala Spark Streaming and Kafka integration. Based on your environment details (Scala 2.11, SBT 1.0, Spark 2.0.1), here are the most common fixes to try:
1. Verify Dependency Consistency in build.sbt
First, make sure your SBT dependencies match exactly the versions you're using in the spark-submit command. Your build.sbt should include:
libraryDependencies ++= Seq( "org.apache.spark" %% "spark-streaming" % "2.0.1" % Provided, "org.apache.spark" %% "spark-streaming-kafka-0-10" % "2.0.1" % Provided )
- The
%%operator automatically aligns the dependency with your Scala 2.11 version, so you don't have to manually append_2.11(which avoids typos). - Using
Providedtells SBT not to package these dependencies into your jar, since Spark's cluster already includes core Spark libraries—this prevents version conflicts.
2. Double-Check the spark-submit --packages Parameter
Your current --packages argument (org.apache.spark:spark-streaming-kafka-0-10_2.11:2.0.1) is correctly formatted, but confirm:
- Your cluster has access to Maven Central Repository to download this dependency. If your cluster is offline, use the
--jarsflag to pass the local Kafka Streaming jar file instead, or pre-install the dependency on all cluster nodes' Spark classpath. - There are no typos in the group ID, artifact ID, or version number (a missing underscore or wrong version is a common culprit).
3. Confirm the Main Class Path is Correct
The --class "KafkaWordCount" argument needs to point to the fully qualified class name. If your KafkaWordCount class is inside a package (e.g., com.example), you must include the package path like com.example.KafkaWordCount.
- To verify, run this command to check if your jar contains the correct class file:
Look for a line likejar tf jars/sskafka_2.11-0.1.jarKafkaWordCount.class(orcom/example/KafkaWordCount.classif using a package).
4. Ensure SBT Packages Your Code Properly
- If you're using
sbt package, your jar will only contain your code (since dependencies are markedProvided), so the--packagesflag inspark-submitmust successfully pull in the required libraries. - If you're using
sbt assemblyto build a fat jar, configure the assembly plugin to exclude Spark core dependencies to avoid conflicts with the cluster's Spark version. Add this to yourbuild.sbt:assemblyMergeStrategy in assembly := { case PathList("META-INF", xs @ _*) => MergeStrategy.discard case x => MergeStrategy.first }
5. Check Cluster Node Classpath Configuration
If your Spark cluster nodes aren't loading the Kafka Streaming dependency correctly, try adding --driver-class-path and --executor-class-path to your spark-submit command, pointing to the location of the Kafka Streaming jar on your cluster. Note that this is usually unnecessary if --packages works, but it can help if your cluster has network or classpath restrictions.
Start with verifying your jar's contents and build.sbt dependencies—those are the most frequent causes of this exception.
内容的提问来源于stack exchange,提问作者Swathi S G

