Azure Databricks升级后读取Apache Iceberg表遇ByteBuffer不支持错误
问题分析
升级Databricks Runtime到14.3 LTS后,Apache Iceberg表可正常写入但读取时报错UnsupportedOperationException: Byte-buffer read unsupported by org.apache.hadoop.fs.BufferedFSInputStream,本质是Iceberg 1.5.2与Azure Hadoop 3.4.0的BufferedFSInputStream存在兼容性冲突——Iceberg读取逻辑依赖ByteBuffer相关方法,但该Hadoop版本的输入流未实现此功能。
解决方案
以下是三种可行的解决思路:
降级Iceberg版本
将Iceberg从1.5.2降级至1.4.3,该版本与Databricks Runtime 14.3 LTS及Azure Hadoop 3.4.0的兼容性经过验证,可规避ByteBuffer读取的兼容问题。- Python环境安装指定版本:
%pip install pyiceberg==0.14.3 - Scala/Spark集群配置:
spark.conf.set("spark.jars.packages", "org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.4.3")
- Python环境安装指定版本:
升级Azure Hadoop依赖
将azure-hadoop版本升级到3.3.5及以上(需匹配Databricks Runtime 14.3的兼容范围),新版本的BufferedFSInputStream已修复ByteBuffer读取不支持的问题。在集群Spark配置中添加:spark.hadoop.fs.azure.enable.append.support true spark.jars.packages org.apache.hadoop:hadoop-azure:3.3.5强制Iceberg使用传统读取方式
通过Spark配置禁用Iceberg的ByteBuffer读取逻辑,改用传统字节数组读取:spark.sql.extensions org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions spark.sql.catalog.spark_catalog org.apache.iceberg.spark.SparkSessionCatalog spark.sql.catalog.spark_catalog.type hive spark.iceberg.hadoop.use-buffer-read false
验证步骤
- 应用任一配置后重启集群
- 执行测试读取代码验证:
df = spark.read.format("iceberg").load("<你的Iceberg表路径>") df.show() - 确认Spark任务正常执行,无报错
内容的提问来源于stack exchange,提问作者Daniel Brenner
相关产品推荐
相关产品推荐

