You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Synapse中如何为嵌套数组结构定义正确的JSON Schema?

正确的JSON Schema定义方案

首先拆解目标JSON中data字段的结构:

  • 顶层是数组,每个元素是包含两个异构元素的数组:
    1. 第一个元素是字符串格式的时间戳(可转为TimestampType)
    2. 第二个元素是二维数值数组(每个子数组包含两个浮点数/整数)

针对这个结构,正确的PySpark Schema定义如下:

from pyspark.sql.types import StructType, StructField, ArrayType, StringType, DoubleType, TimestampType, UnionType

json_schema = StructType([
    StructField("data", ArrayType(
        ArrayType(
            UnionType([
                TimestampType(),  # 若时间戳解析异常,可替换为StringType
                ArrayType(ArrayType(DoubleType(), containsNull=True), containsNull=True)
            ]),
            containsNull=True
        ),
        containsNull=True
    ), nullable=True),
    StructField("indicators", ArrayType(StringType(), containsNull=True), nullable=True),
    StructField("objects", ArrayType(StringType(), containsNull=True), nullable=True)
])

为什么之前的Schema失效?

  • 你第一次定义的四层嵌套数组完全不符合实际结构:data是两层数组,第二层数组内是「字符串+二维数组」的异构组合,而非四层字符串数组,导致解析失败,data列全为null。
  • 自动推断的Schema将data统一转为字符串数组,是因为Spark自动推断时会把数组内的异构类型统一转为字符串,这会导致后续无法对嵌套的数值数组执行explode操作。

额外说明

如果使用TimestampType时出现时间戳解析错误,可先改用StringType读取,之后再通过to_timestamp函数转换格式,避免解析失败。

内容的提问来源于stack exchange,提问作者LYNllow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 19:23:22