You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分DataFrame的JSON数组列并解决类型不匹配问题?

展开DataFrame中的JSON数组列

原始数据

ID   input_array
1    [ {“A”:300, “B”:400}, { “A”:500,”B”: 600} ]
2    [ {“A”: 800, “B”: 900} ]

目标输出

ID A      B
1  300    400
1  500    600
2  800    900

解决方案

你遇到的类型不匹配问题,大概率是两个原因:一是input_array列的内容用了中文引号,JSON无法识别;二是没有正确指定Schema就用from_json解析。按以下步骤处理:

步骤1:统一引号格式(如果是字符串类型的列)

如果input_array是字符串列,先把中文引号替换成英文引号:

df = df.withColumn("input_array", regexp_replace("input_array", "“|”", "\""))

步骤2:定义JSON Schema

提前定义好数组内结构体的Schema,避免解析错误:

from pyspark.sql.types import StructType, StructField, IntegerType, ArrayType

schema = ArrayType(
    StructType([
        StructField("A", IntegerType(), nullable=True),
        StructField("B", IntegerType(), nullable=True)
    ])
)

步骤3:解析JSON并展开

先用from_json把字符串转成数组结构体,再用explode展开数组,最后把结构体的字段拆成列:

from pyspark.sql.functions import from_json, explode

# 解析JSON数组
df_parsed = df.withColumn("parsed_array", from_json("input_array", schema))
# 展开数组
df_exploded = df_parsed.select("ID", explode("parsed_array").alias("struct_col"))
# 拆分结构体字段
final_df = df_exploded.select("ID", "struct_col.A", "struct_col.B")

final_df.show()

如果input_array已经是数组类型(非字符串)

如果你的input_array本身就是数组结构体类型,直接跳过步骤1,执行以下代码即可:

from pyspark.sql.functions import explode

df_exploded = df.select("ID", explode("input_array").alias("struct_col"))
final_df = df_exploded.select("ID", "struct_col.A", "struct_col.B")

内容的提问来源于stack exchange,提问作者Sam Harris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 09:30:45