如何修复Spark运行时InaccessibleObjectException异常
Spark读取文件抛出InaccessibleObjectException异常解决
问题现象
创建本地模式的Spark Session读取文本文件时,程序抛出java.lang.reflect.InaccessibleObjectException,核心错误提示为无法访问java.net.URI类的私有transient字段scheme,原因是java.base模块未向未命名模块@bcec361开放java.net包的反射访问权限。
完整错误栈:
Exception in thread "main" java.lang.reflect.InaccessibleObjectException: Unable to make field private transient java.lang.String java.net.URI.scheme accessible: module java.base does not "opens java.net" to unnamed module @bcec361 at java.base/java.lang.reflect.AccessibleObject.checkCanSetAccessible(AccessibleObject.java:354) at java.base/java.lang.reflect.AccessibleObject.checkCanSetAccessible(AccessibleObject.java:297) at java.base/java.lang.reflect.Field.checkCanSetAccessible(Field.java:178) at java.base/java.lang.reflect.Field.setAccessible(Field.java:172) at org.apache.spark.util.SizeEstimator$$anonfun$getClassInfo$3.apply(SizeEstimator.scala:336) at org.apache.spark.util.SizeEstimator$$anonfun$getClassInfo$3.apply(SizeEstimator.scala:330) at scala.collection.IndexedSeqOptimized$class.foreach(IndexedSeqOptimized.scala:33) at scala.collection.mutable.ArrayOps$ofRef.foreach(ArrayOps.scala:186) at org.apache.spark.util.SizeEstimator$.getClassInfo(SizeEstimator.scala:330) at org.apache.spark.util.SizeEstimator$.visitSingleObject(SizeEstimator.scala:222) at org.apache.spark.util.SizeEstimator$.org$apache$spark$util$SizeEstimator$$estimate(SizeEstimator.scala:201) at org.apache.spark.util.SizeEstimator$.estimate(SizeEstimator.scala:69) at org.apache.spark.sql.execution.datasources.SharedInMemoryCache$$anon$1.weigh(FileStatusCache.scala:109) at org.apache.spark.sql.execution.datasources.SharedInMemoryCache$$anon$1.weigh(FileStatusCache.scala:107) at org.spark_project.guava.cache.LocalCache$Segment.setValue(LocalCache.java:2222) at org.spark_project.guava.cache.LocalCache$Segment.put(LocalCache.java:2944) at org.spark_project.guava.cache.LocalCache.put(LocalCache.java:4212) at org.spark_project.guava.cache.LocalCache$LocalManualCache.put(LocalCache.java:4804) at org.apache.spark.sql.execution.datasources.SharedInMemoryCache$$anon$3.putLeafFiles(FileStatusCache.scala:152) at org.apache.spark.sql.execution.datasources.InMemoryFileIndex$$anonfun$listLeafFiles$2.apply(InMemoryFileIndex.scala:130) at org.apache.spark.sql.execution.datasources.InMemoryFileIndex$$anonfun$listLeafFiles$2.apply(InMemoryFileIndex.scala:128) at scala.collection.mutable.ResizableArray$class.foreach(ResizableArray.scala:59) at scala.collection.mutable.ArrayBuffer.foreach(ArrayBuffer.scala:48) at org.apache.spark.sql.execution.datasources.InMemoryFileIndex.listLeafFiles(InMemoryFileIndex.scala:128) at org.apache.spark.sql.execution.datasources.InMemoryFileIndex.refresh0(InMemoryFileIndex.scala:91) at org.apache.spark.sql.execution.datasources.InMemoryFileIndex.<init>(InMemoryFileIndex.scala:67) at org.apache.spark.sql.execution.datasources.DataSource.org$apache$spark$sql$execution$datasources$DataSource$$createInMemoryFileIndex(DataSource.scala:533) at org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:371) at org.apache.spark.sql.DataFrameReader.loadV1Source(DataFrameReader.scala:223) at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:211) at org.apache.spark.sql.DataFrameReader.text(DataFrameReader.scala:714) at org.apache.spark.sql.DataFrameReader.text(DataFrameReader.scala:686) at org.example.SparkSessionTest$.main(SparkSessionTest.scala:19) at org.example.SparkSessionTest.main(SparkSessionTest.scala)
运行环境
- JDK版本:11.0.15.1
- Scala版本:2.12.10
- Spark版本:3.1.3
- 构建工具:Maven
- 开发工具:IntelliJ IDEA
复现代码
import org.apache.spark.sql.SparkSession object SparkSessionTest { def main(args:Array[String]): Unit ={ val spark = SparkSession.builder() .master("local[1]") .appName("SparkByExample") .getOrCreate(); println("First SparkContext:") println("APP Name :"+spark.sparkContext.appName); println("Deploy Mode :"+spark.sparkContext.deployMode); println("Master :"+spark.sparkContext.master); val df = spark.read.text("src/data/test.txt") } }
问题根因
该异常是JDK9+模块化机制与旧版本Spark反射逻辑不兼容导致。Spark 3.1.x内置的SizeEstimator内存估算组件会通过反射遍历所有加载对象的字段(包括JDK核心类的私有字段)计算内存占用,JDK9引入的模块系统默认禁止未命名模块(业务代码、第三方依赖所在的无模块声明类路径)反射访问java.base核心模块下的私有成员,直接触发访问权限异常。JDK8没有模块化限制,不会出现该问题。
解决方案
方案1:添加JVM启动参数(推荐,无需修改代码和依赖)
启动程序时通过--add-opens参数显式开放对应模块的反射权限,针对本次报错的最小可用参数为:
--add-opens=java.base/java.net=ALL-UNNAMED
为避免后续运行Spark其他逻辑时触发同类反射报错,可一次性添加JDK11下运行Spark 3.1.x所需的全部模块开放参数:
--add-opens=java.base/java.lang=ALL-UNNAMED --add-opens=java.base/java.lang.invoke=ALL-UNNAMED --add-opens=java.base/java.lang.reflect=ALL-UNNAMED --add-opens=java.base/java.io=ALL-UNNAMED --add-opens=java.base/java.net=ALL-UNNAMED --add-opens=java.base/java.nio=ALL-UNNAMED --add-opens=java.base/java.util=ALL-UNNAMED --add-opens=java.base/java.util.concurrent=ALL-UNNAMED --add-opens=java.base/java.util.concurrent.atomic=ALL-UNNAMED --add-opens=java.base/sun.nio.ch=ALL-UNNAMED --add-opens=java.base/sun.nio.cs=ALL-UNNAMED --add-opens=java.base/sun.security.action=ALL-UNNAMED --add-opens=java.base/sun.util.calendar=ALL-UNNAMED
- IDEA配置路径:打开顶部菜单栏
Run > Edit Configurations,找到运行SparkSessionTest的对应配置项,在VM options栏粘贴上述参数,保存后重新运行即可。 - Maven命令行配置:如果通过
mvn exec:java启动程序,可在pom.xml的exec-maven-plugin插件配置中添加上述JVM参数,或在执行命令时通过-Dexec.jvmArgs传递参数。
方案2:切换JDK版本至1.8
Spark 3.1.x系列对JDK8的兼容性最完善,无模块化反射限制。如果项目没有强制JDK版本要求,直接将项目SDK、编译目标版本切换为JDK8即可直接运行,无需额外配置。
注意:Spark 3.1.x不支持JDK17及以上版本,强行使用会触发更多兼容性问题,高版本JDK需要升级Spark到3.3+版本才能获得官方支持。
内容的提问来源于stack exchange,提问作者soao
相关产品推荐
相关产品推荐

