You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark 2.4中使用Avro内置数据源运行代码报错求助

Fixing "Failed to find data source: avro" in Spark 2.4

Hey there, let's get this Avro data source issue sorted out! The error message you're seeing is pretty clear—starting with Spark 2.4, Avro is no longer included in the core Spark distribution by default. It's a built-in but external module, which means you need to explicitly include its dependency when running your Spark code.

Here are the most common solutions depending on how you're running your Spark application:

1. Running in Spark Shell

When launching the Spark Shell, add the Avro package using the --packages flag. Make sure the Scala version suffix (e.g., _2.11) and package version match your Spark installation:

spark-shell --packages org.apache.spark:spark-avro_2.11:2.4.8

Note: Replace 2.11 with your Spark's Scala version (most Spark 2.4 distributions use 2.11) and 2.4.8 with your exact Spark version number.

2. Using SBT for Project Builds

Add the Avro dependency to your build.sbt file. The Provided scope tells SBT not to package the dependency with your JAR (since it'll be available on the Spark cluster):

libraryDependencies += "org.apache.spark" %% "spark-avro" % "2.4.8" % Provided

If you're testing locally (not on a cluster), you can remove the % Provided part so the dependency is included in your classpath.

3. Using Maven for Project Builds

Add the following dependency to your pom.xml file. Again, match the Scala version and Spark version to your setup:

<dependency>
    <groupId>org.apache.spark</groupId>
    <artifactId>spark-avro_2.11</artifactId>
    <version>2.4.8</version>
    <scope>provided</scope>
</dependency>

Like with SBT, remove <scope>provided</scope> if you're running locally.

4. Submitting a Spark Application with spark-submit

Include the Avro package when submitting your JAR using --packages:

spark-submit --packages org.apache.spark:spark-avro_2.11:2.4.8 --com.your.package.MainClass your-app.jar

Key Notes to Avoid Headaches:

  • Version Alignment: Always use the exact same version for spark-avro as your Spark core version. Mismatched versions will lead to compatibility errors.
  • Scala Version: The suffix _2.11 or _2.12 must match the Scala version your Spark was compiled with. Check your Spark docs or run spark-shell and look at the welcome message to confirm.
  • IDE Local Runs: If you're using an IDE like IntelliJ, make sure the spark-avro dependency is added to your project's build file and that your run configuration includes it in the classpath.

Once you add the correct dependency, your code val usersDF = spark.read.format("avro").load("examples/src/main/resources/users.avro") should run without the AnalysisException error.

内容的提问来源于stack exchange,提问作者Achilleus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:24:59