You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何创建PySpark Schema读取含异构类型数组的JSON文件?

解决Spark Schema中混合类型数组的定义问题

你的prediction_probability字段是一个混合类型数组,元素交替为[整数, 整数]格式的数组和浮点数,Spark中需要用UnionType(联合类型)来定义这种包含多种类型元素的数组。

具体Schema定义

首先导入必要的Spark数据类型:

from pyspark.sql.types import StructType, StructField, ArrayType, IntegerType, DoubleType, UnionType

然后定义该字段的Schema:

schema = StructType([
    StructField(
        name="prediction_probability",
        dataType=ArrayType(
            UnionType([
                ArrayType(IntegerType()),  # 对应[0,0]这类整数数组
                DoubleType()               # 对应0.0788这类浮点数
            ])
        ),
        nullable=True
    )
])

说明

  • UnionType用于声明元素可以是多种类型中的一种,这里指定了两种可能的元素类型:包含两个整数的数组,以及双精度浮点数。
  • 这样定义后,Spark就能正确解析[[0,0],0.0788,[1,0],0.0015]这类混合结构的数组字段。

内容的提问来源于stack exchange,提问作者Bjarne Pedersen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 19:10:33