自定义Spark Schema适配JSON报错,该如何修复?
问题描述
加载如下JSON数据到Spark DataFrame(未指定Schema):
{ "titles": { "L": [ { "S": "ABC" } ] } }
执行df.printSchema()后自动推断的Schema为:
root |-- titles: struct (nullable = true) | |-- L: array (nullable = true) | | |-- element: struct (containsNull = true) | | | |-- S: string (nullable = true)
尝试手动定义如下Schema读取同一JSON时失败:
AS = StructType([StructField ("L", ArrayType(StructField("S", StringType(), True)) ) ]) my_schema = StructType([ StructField("titles", AS ,True) ])
报错信息:
Failed to convert the JSON string '{"metadata":{},"name":"S","nullable":true,"type":"string"}' to a data type
修复方案
错误根源是ArrayType的参数必须是完整的StructType对象,而非直接传入StructField。需要将数组内的元素结构用StructType包裹。
正确的Schema定义代码如下:
from pyspark.sql.types import StructType, StructField, ArrayType, StringType # 定义数组元素的结构体 element_struct = StructType([ StructField("S", StringType(), nullable=True) ]) # 定义titles字段的结构体 titles_struct = StructType([ StructField("L", ArrayType(element_struct), nullable=True) ]) # 最终完整Schema my_schema = StructType([ StructField("titles", titles_struct, nullable=True) ])
也可以使用简化的嵌套写法:
my_schema = StructType([ StructField("titles", StructType([ StructField("L", ArrayType(StructType([ StructField("S", StringType(), nullable=True) ])), nullable=True) ]), nullable=True) ])
使用上述定义的Schema读取JSON数据即可正常执行。
内容的提问来源于stack exchange,提问作者user3440012
相关产品推荐
相关产品推荐

