You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Streaming中如何定义Schema解析包含对象数组的JSON数据

Spark 嵌套对象数组字段Schema定义方案

对于JSON结构中的对象数组类型字段,在Spark中使用ArrayType嵌套对应元素的StructType即可完成定义,无需额外做explode操作,可直接匹配BigQuery的REPEATED RECORD类型。

完整Schema定义代码

import org.apache.spark.sql.types.{StructType, StringType, ArrayType}

val tableSchema: StructType = (new StructType)
  .add("messageDetails", (new StructType)
    .add("id", StringType)
    .add("name", StringType))
  .add("messageMain", (new StructType)
    .add("date", StringType)
    // details为数组类型,数组内每个元素是包含val1、val2的结构体
    .add("details", ArrayType(
      new StructType()
        .add("val1", StringType)
        .add("val2", StringType)
    ))
  )

写入BigQuery说明

按上述方式定义Schema读取JSON数据后,messageMain.details字段在Spark中为Array[Struct(val1: String, val2: String)]类型,使用Spark BigQuery连接器写入时,会自动映射为BigQuery表要求的REPEATED RECORD类型,全程不需要对数组做展开处理,完全符合使用要求。

内容的提问来源于stack exchange,提问作者asdasd32

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 01:06:04