Scala读取含空Struct列的JSON文件生成DataFrame失败问题
问题原因
你的JSON文件中parent字段同时存在两种类型:空字符串(String)和结构化对象(Struct)。Spark自动推断Schema时,会将该字段统一识别为String类型,结构化对象会被序列化为JSON字符串。后续执行select(s"$parentNode.*")后,尝试访问parent的子字段时会失败,因为String类型没有嵌套字段。
解决方案
通过以下步骤统一parent字段的类型,将空字符串转为null,结构化字符串解析为Struct类型:
- 定义
parent字段的Struct Schema - 读取JSON时指定Schema(可选,但能避免自动推断的问题)
- 转换
parent字段类型,统一为StructType
完整Scala代码示例:
import org.apache.spark.sql.types._ import org.apache.spark.sql.functions._ // 定义parent字段的结构化Schema val parentSchema = StructType(Seq( StructField("display_value", StringType, nullable = true), StructField("link", StringType, nullable = true) )) // 定义整个JSON的完整Schema(可选,用于强制指定字段类型) val fullSchema = StructType(Seq( StructField("result", ArrayType(StructType(Seq( StructField("parent", StringType, nullable = true), StructField("roles", StringType, nullable = true), StructField("sys_created_by", StringType, nullable = true) ))), nullable = true) )) // 读取JSON文件,指定Schema避免自动推断的类型冲突 var df: DataFrame = spark.read .option("multiline", "true") .option("mode", "PERMISSIVE") .schema(fullSchema) // 若不指定,Spark会自动推断为String,但指定更可靠 .json(path) // 处理数据并转换parent字段类型 var updatedDf: DataFrame = { if (null != parentNode && parentNode.trim.nonEmpty) { df.select(explode(col(parentNode)).as(parentNode)) .select(s"$parentNode.*") // 将空字符串转为null,非空字符串解析为Struct .withColumn("parent", when(col("parent") === "", lit(null)).otherwise(from_json(col("parent"), parentSchema))) } else { df .withColumn("parent", when(col("parent") === "", lit(null)).otherwise(from_json(col("parent"), parentSchema))) } } // 验证结果 updatedDf.printSchema() updatedDf.show(truncate = false)
说明
- 指定Schema可以确保Spark按预期读取字段类型,避免自动推断带来的类型冲突
from_json函数将JSON字符串解析为StructType,when表达式处理空字符串的情况,转为null保持类型统一- 处理后
parent字段为StructType,空值记录的parent为null,可正常访问其嵌套字段(如parent.display_value)
内容的提问来源于stack exchange,提问作者Pooja Upadhyay
相关产品推荐
相关产品推荐

