You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义Spark Schema适配JSON报错,该如何修复?

问题描述

加载如下JSON数据到Spark DataFrame(未指定Schema):

{
 "titles": {
  "L": [
   {
    "S": "ABC"
   }
  ]
 }
}

执行df.printSchema()后自动推断的Schema为:

root
 |-- titles: struct (nullable = true)
 |    |-- L: array (nullable = true)
 |    |    |-- element: struct (containsNull = true)
 |    |    |    |-- S: string (nullable = true)

尝试手动定义如下Schema读取同一JSON时失败:

AS = StructType([StructField
  ("L",
    ArrayType(StructField("S", StringType(), True))
 )   
]) 

my_schema = StructType([
   StructField("titles", AS ,True)
])

报错信息:

Failed to convert the JSON string '{"metadata":{},"name":"S","nullable":true,"type":"string"}' to a data type
修复方案

错误根源是ArrayType的参数必须是完整的StructType对象,而非直接传入StructField。需要将数组内的元素结构用StructType包裹。

正确的Schema定义代码如下:

from pyspark.sql.types import StructType, StructField, ArrayType, StringType

# 定义数组元素的结构体
element_struct = StructType([
    StructField("S", StringType(), nullable=True)
])

# 定义titles字段的结构体
titles_struct = StructType([
    StructField("L", ArrayType(element_struct), nullable=True)
])

# 最终完整Schema
my_schema = StructType([
    StructField("titles", titles_struct, nullable=True)
])

也可以使用简化的嵌套写法:

my_schema = StructType([
    StructField("titles", StructType([
        StructField("L", ArrayType(StructType([
            StructField("S", StringType(), nullable=True)
        ])), nullable=True)
    ]), nullable=True)
])

使用上述定义的Schema读取JSON数据即可正常执行。

内容的提问来源于stack exchange,提问作者user3440012

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 01:25:25