You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在AWS Glue中将并行数组合并为DataFrame结构

AWS Glue 实现并行数组关联转换DataFrame

需求说明

原始JSON结构:

{
    "out": [
        {
            "attr": [ "a1", "a2", "a3" ],
            "val": [ 1, 2, 3 ],
            "text": "test1"
        },
        {
            "attr": [ "a4", "a5", "a6" ],
            "val": [ 4, 5, 6 ],
            "text": "test2"
        }
    ],
    "ids": [
        "id1",
        "id2"
    ]    
}

其中ids与out是并行数组,期望转换为如下结构的DataFrame:

id     text     attr            val
--     ----     ----            ---
id1    test1    [a1, a2, a3]    [1,2,3]
id2    test2    [a4, a5, a6]    [4,5,6]

解决方案

核心思路是利用posexplode函数展开数组时保留位置索引,通过索引关联两个并行数组的对应元素,实现横向拼接。

完整代码示例

from pyspark.sql.functions import posexplode

spark_context = SparkContext.getOrCreate()
glue_context = GlueContext(spark_context)
spark = glue_context.spark_session

# 加载JSON数据
df = spark.read.json("<location>")

# 展开ids数组,保留位置索引
df_ids = df.select(posexplode("ids").alias("idx", "id"))

# 展开out数组,保留位置索引,并解析结构体字段
df_out = df.select(posexplode("out").alias("idx", "out_struct")) \
           .select("idx", "out_struct.text", "out_struct.attr", "out_struct.val")

# 通过索引关联两个DataFrame,删除冗余索引列
result_df = df_ids.join(df_out, on="idx", how="inner").drop("idx")

# 查看结果
result_df.show()

关键说明

  • posexplode:将数组的每个元素与其对应的位置索引(从0开始)一起展开,确保并行数组的对应元素拥有相同的idx值。
  • 关联逻辑:通过idx字段将两个展开后的DataFrame做内连接,即可得到一一对应的最终结果。
  • 注意:若你的DataFrame中数组字段名是id(而非原始JSON中的ids,以实际Schema为准),需将代码中的posexplode("ids")改为posexplode("id")。

内容的提问来源于stack exchange,提问作者Jaco Van Niekerk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 19:02:32