如何用PySpark快速展开深度嵌套的Struct结构?
PySpark展开嵌套结构化字段的方法
问题描述
示例数据集
+---+------------------------------+ |id |example_field | +---+------------------------------+ |1 |{[{[{111, AAA}, {222, BBB}]}]}| +---+------------------------------+
字段数据类型
[('id', 'int'), ('example_field', 'struct<xxx:array<struct<nested_field:array<struct<field_1:int,field_2:string>>>>>')]
期望输出:
id field_1 field_2 1 111 AAA 1 222 BBB
解决方案
可以通过PySpark的explode函数逐层展开嵌套的数组和结构体,具体步骤如下:
- 展开最外层数组字段
xxx:
from pyspark.sql.functions import explode df = df.withColumn("exploded_xxx", explode("example_field.xxx"))
- 展开第二层数组字段
nested_field:
df = df.withColumn("exploded_nested", explode("exploded_xxx.nested_field"))
- 提取目标字段并清理中间字段:
result_df = df.select( "id", "exploded_nested.field_1", "exploded_nested.field_2" )
也可以合并为链式调用,简化代码:
from pyspark.sql.functions import explode result_df = df \ .withColumn("exploded_xxx", explode("example_field.xxx")) \ .withColumn("exploded_nested", explode("exploded_xxx.nested_field")) \ .select("id", "exploded_nested.field_1", "exploded_nested.field_2")
执行上述代码后,即可得到你需要的扁平结构数据集。
内容的提问来源于stack exchange,提问作者wawawa
相关产品推荐
相关产品推荐

