You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Structured Streaming连接Kafka报错"Failed to find data source: kafka"求助

Troubleshooting ClassNotFoundException: Failed to find data source: kafka in Spark Structured Streaming

Hey there, let's work through why you're hitting this error even though you've added dependencies and confirmed Kafka/ZooKeeper are up and running. Here are the most common fixes to check:

1. Verify Dependency Version Compatibility

This is the #1 culprit. The Spark Kafka connector version must match your Spark version exactly, and you also need to align the Scala version suffix (like _2.12) with the one your Spark distribution uses.

For example, if you're using Spark 3.5.0 with Scala 2.12, your dependency should look like this:

  • Maven:
    <dependency>
        <groupId>org.apache.spark</groupId>
        <artifactId>spark-sql-kafka-0-10_2.12</artifactId>
        <version>3.5.0</version>
    </dependency>
    
  • Gradle:
    implementation 'org.apache.spark:spark-sql-kafka-0-10_2.12:3.5.0'
    

Mismatched versions (e.g., using a 3.2.x connector with Spark 3.5.x) will almost always cause classpath issues.

2. Check Dependency Scope & Deployment Method

If you're using spark-submit to run your app, adding the dependency to your project's build file might not be enough—you need to ensure the connector is included in the runtime classpath.

Instead of relying on local build dependencies, use the --packages flag with spark-submit to pull the correct connector dynamically:

spark-submit --packages org.apache.spark:spark-sql-kafka-0-10_2.12:3.5.0 your-application.jar

If you're using a build tool like Maven/Gradle, avoid setting the dependency scope to provided unless you're sure the cluster already has the connector jar (most don't). Use compile (Maven) or implementation (Gradle) to include it in your build.

3. Validate Classpath Configuration

  • Local IDE Run: Double-check that your IDE (IntelliJ/Eclipse) has loaded all dependencies correctly. Sometimes cached libraries can cause issues—try refreshing your project's dependencies or invalidating caches.
  • Cluster Mode: If running on a cluster, ensure the spark.driver.extraClassPath and spark.executor.extraClassPath properties include the path to the Spark Kafka connector jar. Alternatively, place the jar in Spark's jars directory on all nodes (though this is less flexible than using --packages).

4. Ensure You're Using the Correct Dependency Artifact

Don't confuse the plain Kafka client library (kafka-clients) with the Spark-specific Kafka connector. The latter is what provides the kafka data source for Structured Streaming. Make sure your build file includes spark-sql-kafka-0-10 (not just kafka-clients).

5. Confirm Spark Session Code is Correct

Double-check that you're specifying the data source correctly in your code. The format should be exactly "kafka", not a variant like "kafka-0-10":

val streamingDF = spark.readStream
  .format("kafka") // This is critical—don't misspell or use an old format name
  .option("kafka.bootstrap.servers", "localhost:9092")
  .option("subscribe", "your-target-topic")
  .load()

If you've gone through all these steps and still see the error, sharing your exact Spark version, build configuration (pom.xml/build.sbt snippet), and deployment method would help narrow things down further.

内容的提问来源于stack exchange,提问作者Rahul Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:41:54