使用Databricks Autoloader读取ADLS Gen2 JSON文件遇类型不匹配错误
问题
在Delta Live Pipeline中使用Databricks Autoloader读取Azure ADLS Gen2中的JSON文件时,仅部分特定文件出现报错,已确认这些文件未损坏。
核心报错信息:
java.lang.IllegalArgumentException: requirement failed: Literal must have a corresponding value to string, but class Integer found
完整错误栈:
Caused by: java.lang.IllegalArgumentException: requirement failed: Literal must have a corresponding value to string, but class Integer found. at scala.Predef$.require(Predef.scala:281) at com.databricks.sql.io.FileReadException: Error while reading file /mnt/Source/kafka/customer_raw/filtered_data/year=2022/month=11/day=9/hour=15/part-00000-31413bcf-0a8f-480f-8d45-6970f4c4c9f7.c000.json. at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1$$anon$2.logFileNameAndThrow(FileScanRDD.scala:598) at org.apache.spark.sql.execution.datasources.FileScanRDD$$anon$1.hasNext(FileScanRDD.scala:422) at scala.collection.Iterator$$anon$10.hasNext(Iterator.scala:460) at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIteratorForCodegenStage1.processNext(null:-1) at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43) at org.apache.spark.sql.execution.WholeStageCodegenExec$$anon$1.hasNext(WholeStageCodegenExec.scala:759) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) at java.lang.Thread.run(Thread.java:750) java.lang.IllegalArgumentException: requirement failed: Literal must have a corresponding value to string, but class Integer found. at scala.Predef$.require(Predef.scala:281) at org.apache.spark.sql.catalyst.expressions.Literal$.validateLiteralValue(literals.scala:274) org.apache.spark.sql.execution.WholeStageCodegenExec$$anon$1.hasNext(WholeStageCodegenExec.scala) at java.lang.Thread.run(Thread.java:750)
使用的Delta Live Pipeline代码(原代码存在语法错误,已标注修正点):
@dlt.table(name = tablename, comment = "Create Bronze Table", table_properties={ "quality": "bronze" } ) def Bronze_Table_Create(): return ( spark .readStream .schema(schemapath) .format("cloudFiles") .option("cloudFiles.format","json") # 原代码缺失闭合引号 .option("cloudFiles.schemaLocation", schemalocation) # 原代码参数名拼写错误+缺失引号 .option("cloudFiles.inferColumnTypes", "false") .option("cloudFiles.schemaEvolutionMode", "rescue") .load(sourcelocation) # 原代码缺失闭合括号 )
解决方案
- 字段类型不匹配处理:预定义schema中某字段被设为String类型,但报错文件内对应字段实际是Integer类型。虽然开启了
rescue模式,但可能因schema加载逻辑异常未触发字段救援。- 定位冲突字段:对比预定义schema与报错文件的具体内容,找出类型不匹配的字段
- 调整字段定义:将schema中冲突字段改为Integer类型;若需保留String类型,可在读取后添加转换逻辑,如
cast(col("field_name").cast(StringType));或调整schemaEvolutionMode为addNewColumns(需确保现有字段类型兼容)
- 修正代码语法错误:原代码的引号缺失、参数名拼写错误会导致Autoloader的schema处理逻辑异常,必须修正所有语法问题
- 检查版本兼容性:部分Databricks Runtime版本存在
rescue模式处理类型不匹配字段的bug,确认当前DLT对应Runtime版本是否支持配置,必要时升级版本 - 临时验证无预定义schema场景:移除预定义schema,让Autoloader自动推断schema,验证是否仍报错,再根据推断结果调整最终schema
内容的提问来源于stack exchange,提问作者swapnil kamle
相关产品推荐
相关产品推荐

