You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PySpark快速展开深度嵌套的Struct结构?

PySpark展开嵌套结构化字段的方法

问题描述

示例数据集

+---+------------------------------+
|id |example_field                  |
+---+------------------------------+
|1  |{[{[{111, AAA}, {222, BBB}]}]}|
+---+------------------------------+

字段数据类型

[('id', 'int'),
 ('example_field',
  'struct<xxx:array<struct<nested_field:array<struct<field_1:int,field_2:string>>>>>')]

期望输出:

id  field_1 field_2
1   111     AAA
1   222     BBB

解决方案

可以通过PySpark的explode函数逐层展开嵌套的数组和结构体,具体步骤如下:

  1. 展开最外层数组字段xxx:
from pyspark.sql.functions import explode

df = df.withColumn("exploded_xxx", explode("example_field.xxx"))
  1. 展开第二层数组字段nested_field:
df = df.withColumn("exploded_nested", explode("exploded_xxx.nested_field"))
  1. 提取目标字段并清理中间字段:
result_df = df.select(
    "id",
    "exploded_nested.field_1",
    "exploded_nested.field_2"
)

也可以合并为链式调用,简化代码:

from pyspark.sql.functions import explode

result_df = df \
    .withColumn("exploded_xxx", explode("example_field.xxx")) \
    .withColumn("exploded_nested", explode("exploded_xxx.nested_field")) \
    .select("id", "exploded_nested.field_1", "exploded_nested.field_2")

执行上述代码后,即可得到你需要的扁平结构数据集。

内容的提问来源于stack exchange,提问作者wawawa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 20:10:32