You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark DataFrame中读取复杂JSON数组?

解决PySpark读取嵌套数组JSON无数据显示问题

你的JSON是二维嵌套数组结构(外层数组包含多个内层数组,每个内层数组是对象集合),直接用multiLine读取时,Spark会把整个外层数组解析成一个名为value的数组列,而非直接展开为多行数据,所以你看到的只有列结构没有具体数据行。

解决步骤:

需要通过两次explode函数展开嵌套数组,将二维数组扁平化:

from pyspark.sql.functions import explode

# 读取JSON文件
df = spark.read.format("json") \
    .option("inferSchema", "true") \
    .option("multiLine", "true") \
    .load("/mnt/blob/input/jsonfile.json")

# 第一步:展开外层数组,得到每个内层数组
df_outer = df.select(explode("value").alias("inner_array"))

# 第二步:展开内层数组,得到单个对象行,再提取对象字段
final_df = df_outer.select(explode("inner_array").alias("item")).select("item.*")

# 查看最终结果
final_df.show()

可选:指定Schema(推荐)

如果JSON结构固定,建议手动定义Schema替代inferSchema,提升性能并避免类型推断错误:

from pyspark.sql.types import StructType, StructField, StringType, ArrayType

# 定义匹配JSON结构的Schema
json_schema = ArrayType(
    ArrayType(
        StructType([StructField("Key", StringType(), nullable=True)])
    )
)

# 读取数据
df = spark.read.format("json") \
    .schema(json_schema) \
    .option("multiLine", "true") \
    .load("/mnt/blob/input/jsonfile.json")

# 后续展开操作同上

内容的提问来源于stack exchange,提问作者Luukv93

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 16:50:25