如何在AWS Glue中将并行数组合并为DataFrame结构
AWS Glue 实现并行数组关联转换DataFrame
需求说明
原始JSON结构:
{ "out": [ { "attr": [ "a1", "a2", "a3" ], "val": [ 1, 2, 3 ], "text": "test1" }, { "attr": [ "a4", "a5", "a6" ], "val": [ 4, 5, 6 ], "text": "test2" } ], "ids": [ "id1", "id2" ] }
其中ids与out是并行数组,期望转换为如下结构的DataFrame:
id text attr val -- ---- ---- --- id1 test1 [a1, a2, a3] [1,2,3] id2 test2 [a4, a5, a6] [4,5,6]
解决方案
核心思路是利用posexplode函数展开数组时保留位置索引,通过索引关联两个并行数组的对应元素,实现横向拼接。
完整代码示例
from pyspark.sql.functions import posexplode spark_context = SparkContext.getOrCreate() glue_context = GlueContext(spark_context) spark = glue_context.spark_session # 加载JSON数据 df = spark.read.json("<location>") # 展开ids数组,保留位置索引 df_ids = df.select(posexplode("ids").alias("idx", "id")) # 展开out数组,保留位置索引,并解析结构体字段 df_out = df.select(posexplode("out").alias("idx", "out_struct")) \ .select("idx", "out_struct.text", "out_struct.attr", "out_struct.val") # 通过索引关联两个DataFrame,删除冗余索引列 result_df = df_ids.join(df_out, on="idx", how="inner").drop("idx") # 查看结果 result_df.show()
关键说明
posexplode:将数组的每个元素与其对应的位置索引(从0开始)一起展开,确保并行数组的对应元素拥有相同的idx值。- 关联逻辑:通过
idx字段将两个展开后的DataFrame做内连接,即可得到一一对应的最终结果。 - 注意:若你的DataFrame中数组字段名是
id(而非原始JSON中的ids,以实际Schema为准),需将代码中的posexplode("ids")改为posexplode("id")。
内容的提问来源于stack exchange,提问作者Jaco Van Niekerk
相关产品推荐
相关产品推荐

