You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark嵌套JSON Schema编写及空值问题解决方法

解决嵌套JSON的Spark Schema定义问题

一、确定正确数据类型的方法

  • 直接解析样例JSON:拿出一条典型的原始JSON数据,逐层拆解字段类型。比如gps_coordinates是子结构,里面通常包含latitude(数值型)、longitude(数值型)这类字段,而非字符串。
  • 用Spark自动推断做参考:先用少量测试数据让Spark自动生成Schema,命令如下:
    val df = spark.read.json("/path/to/sample-data.json")
    df.printSchema()
    
    输出的结构会明确每个字段的层级、类型和是否可为空,直接参考这个结果来手动编写Schema即可。

二、编写正确的嵌套Schema步骤

以常见的嵌套JSON结构为例,假设原始数据样例为:

{
  "record_id": "REC001",
  "place_results": {
    "title": "Empire State Building",
    "full_address": "350 5th Ave, New York, NY 10118",
    "gps_coordinates": {
      "lat": 40.7484,
      "lon": -73.9857
    },
    "rating": 4.8
  }
}
  1. 先定义最内层的子结构Schema:也就是gps_coordinates的结构
    import org.apache.spark.sql.types._
    
    val gpsSchema = StructType(Seq(
      StructField("lat", DoubleType, nullable = true),
      StructField("lon", DoubleType, nullable = true)
    ))
    
  2. 定义place_results的Schema:包含它的所有子字段,其中gps_coordinates使用上面定义的子Schema
    val placeResultsSchema = StructType(Seq(
      StructField("title", StringType, nullable = true),
      StructField("full_address", StringType, nullable = true),
      StructField("gps_coordinates", gpsSchema, nullable = true),
      StructField("rating", DoubleType, nullable = true)
    ))
    
  3. 定义完整的顶层Schema:包含所有顶层字段
    val fullSchema = StructType(Seq(
      StructField("record_id", StringType, nullable = true),
      StructField("place_results", placeResultsSchema, nullable = true)
    ))
    

三、踩坑说明

  • 之前把gps_coordinates定义为StringType,但实际是嵌套Struct,类型不匹配会导致Spark无法解析,最终返回空值。
  • 字段缺失是因为编写Schema时没有覆盖place_results里的所有子字段,必须确保每个层级的字段都在Schema中定义。

内容的提问来源于stack exchange,提问作者Xi12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 20:05:37