Spark 2.4中使用Avro内置数据源运行代码报错求助
Hey there, let's get this Avro data source issue sorted out! The error message you're seeing is pretty clear—starting with Spark 2.4, Avro is no longer included in the core Spark distribution by default. It's a built-in but external module, which means you need to explicitly include its dependency when running your Spark code.
Here are the most common solutions depending on how you're running your Spark application:
1. Running in Spark Shell
When launching the Spark Shell, add the Avro package using the --packages flag. Make sure the Scala version suffix (e.g., _2.11) and package version match your Spark installation:
spark-shell --packages org.apache.spark:spark-avro_2.11:2.4.8
Note: Replace 2.11 with your Spark's Scala version (most Spark 2.4 distributions use 2.11) and 2.4.8 with your exact Spark version number.
2. Using SBT for Project Builds
Add the Avro dependency to your build.sbt file. The Provided scope tells SBT not to package the dependency with your JAR (since it'll be available on the Spark cluster):
libraryDependencies += "org.apache.spark" %% "spark-avro" % "2.4.8" % Provided
If you're testing locally (not on a cluster), you can remove the % Provided part so the dependency is included in your classpath.
3. Using Maven for Project Builds
Add the following dependency to your pom.xml file. Again, match the Scala version and Spark version to your setup:
<dependency> <groupId>org.apache.spark</groupId> <artifactId>spark-avro_2.11</artifactId> <version>2.4.8</version> <scope>provided</scope> </dependency>
Like with SBT, remove <scope>provided</scope> if you're running locally.
4. Submitting a Spark Application with spark-submit
Include the Avro package when submitting your JAR using --packages:
spark-submit --packages org.apache.spark:spark-avro_2.11:2.4.8 --com.your.package.MainClass your-app.jar
Key Notes to Avoid Headaches:
- Version Alignment: Always use the exact same version for
spark-avroas your Spark core version. Mismatched versions will lead to compatibility errors. - Scala Version: The suffix
_2.11or_2.12must match the Scala version your Spark was compiled with. Check your Spark docs or runspark-shelland look at the welcome message to confirm. - IDE Local Runs: If you're using an IDE like IntelliJ, make sure the
spark-avrodependency is added to your project's build file and that your run configuration includes it in the classpath.
Once you add the correct dependency, your code val usersDF = spark.read.format("avro").load("examples/src/main/resources/users.avro") should run without the AnalysisException error.
内容的提问来源于stack exchange,提问作者Achilleus

