如何在HDInsight Spark/Jupyter中使用Avro?读取文件报错求助
Hey there, let's get that Avro reading issue sorted out in your HDInsight Spark/Jupyter cluster. The error you're hitting happens because Spark can't locate the required Avro data source library—here are actionable ways to fix it, depending on your use case:
If you just need to work with Avro for a single notebook session, configure your SparkSession to pull the required library on startup. Make sure the package version matches your Spark and Scala versions (most HDInsight Spark 2.4 clusters use Scala 2.11, so we'll use that as an example):
from pyspark.sql import SparkSession spark = SparkSession.builder \ .appName("AvroProcessing") \ .config("spark.jars.packages", "com.databricks:spark-avro_2.11:4.0.0") \ .getOrCreate()
Once the session starts, you can read Avro files like this:
df = spark.read.format("com.databricks.spark.avro").load("/path/to/your/avro/files")
For Spark 3.x+ clusters (HDInsight 4.0+), note that Avro support is built into Spark natively—you don't need the Databricks package anymore. Just use:
df = spark.read.format("avro").load("/path/to/your/avro/files")
If you or your team work with Avro regularly, configuring the dependency at the cluster level avoids adding it to every notebook. Here's how:
- Log into the HDInsight Ambari portal for your cluster.
- Navigate to Spark2 > Configs > Advanced > spark-defaults.
- Add a new property:
spark.jars.packages com.databricks:spark-avro_2.11:4.0.0 - Save the changes and restart the Spark services when prompted. Now every Spark session (including Jupyter) will automatically load the Avro library.
If you prefer using Spark Magic in Jupyter, pass the package dependency directly in the magic command:
For Python:
%spark.pyspark --packages com.databricks:spark-avro_2.11:4.0.0 df = spark.read.format("com.databricks.spark.avro").load("/path/to/avro") df.show()
For Scala:
%%spark --packages com.databricks:spark-avro_2.11:4.0.0 val df = spark.read.format("com.databricks.spark.avro").load("/path/to/avro") df.show()
Quick version check tip
If you're unsure about your Spark/Scala version, run this in your notebook to confirm:
print(f"Spark version: {spark.version}") print(f"Scala version: {spark.sparkContext._jvm.scala.util.Properties.versionString()}")
内容的提问来源于stack exchange,提问作者Jiew Meng

