如何将DataFrame中含嵌套结构的数组列展开为多行?
展开DataFrame数组列为多行的实现方法
问题背景
现有一个DataFrame,其schema定义如下:
root |--sentences:array | |--element:string
该sentences列的元素是包含嵌套字典与列表的数组对象(以下为简化示例):
[{sentences=hello, id=1234, lang=eng, attributes={lang=eng, sid=789, number=[+10000000000]}, {sentences=hello2, id=1234, lang=eng, attributes={lang=eng, sid=789, number=[+10000000000]}]
需要将该数组列展开为多行,每行对应数组中的一个元素。
实现方法
1. Spark DataFrame 场景
使用Spark内置的explode函数即可完成数组列的拆分:
from pyspark.sql.functions import explode # 假设原DataFrame名为df exploded_df = df.select(explode("sentences").alias("sentence"))
如果需要保留原DataFrame的其他列,可在select中同时指定:
exploded_df = df.select("*", explode("sentences").alias("sentence"))
执行后,原数组中的每个元素会单独成为一行,新列sentence存储拆分后的单个对象。
2. Pandas DataFrame 场景
Pandas自带explode方法,直接作用于目标列即可:
# 假设原DataFrame名为df exploded_df = df.explode("sentences", ignore_index=True)
参数ignore_index=True会重置输出结果的行索引,保证索引连续。拆分后每行对应原数组中的一个元素。
内容的提问来源于stack exchange,提问作者ganesh
相关产品推荐
相关产品推荐

