You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在AWS Glue的Spark环境中展开嵌套JSON生成目标DataFrame

问题原因
  • 第一个select("stats.*")报错:stats字段当前被Spark推断为数组类型(ArrayType),只有结构体(StructType)类型才支持.*语法直接展开字段
  • 第二个explode报错:你的stats字段被推断为Map类型,explode处理Map类型会返回key、value两列,仅指定1个别名会和输出列数不匹配
解决方案

方案1:直接处理现有DF(无需修改DF创建逻辑)

如果stats数组中仅包含1个结构体元素,用getItem(0)取出数组元素后再展开即可:

from pyspark.sql.functions import col

df_expanded = df.select(
    "start_time",
    "end_time",
    col("stats").getItem(0).alias("stats_obj")
).select(
    "start_time",
    "end_time",
    "stats_obj.impressions",
    "stats_obj.swipes",
    "stats_obj.view_completion",
    "stats_obj.spend"
)

如果stats数组包含多个元素,用explode先炸开数组再展开:

from pyspark.sql.functions import explode

df_expanded = df.select(
    "start_time",
    "end_time",
    explode("stats").alias("stats_obj")
).select("start_time","end_time","stats_obj.*")

方案2:指定schema创建DF(更稳定,推荐)

Spark自动推断数据类型容易出现偏差,你可以在创建DF时手动指定schema,后续直接用你最初的写法即可:

from pyspark.sql.types import StructType, StructField, StringType, IntegerType, DoubleType

# 定义stats字段的结构体 schema
stats_schema = StructType([
    StructField("impressions", IntegerType(), nullable=True),
    StructField("swipes", IntegerType(), nullable=True),
    StructField("view_completion", DoubleType(), nullable=True),
    StructField("spend", DoubleType(), nullable=True)
])

# 定义单条时间序列数据的schema
timeseries_schema = StructType([
    StructField("start_time", StringType(), nullable=True),
    StructField("end_time", StringType(), nullable=True),
    StructField("stats", stats_schema, nullable=True)
])

# 创建DF时传入指定schema
rdd = sc.parallelize(JSON_resp['timeseries_stats'][0]['timeseries_stat']['timeseries'])
df = rdd.toDF(schema=timeseries_schema)

# 直接展开即可,和你最开始的写法一致
df_expanded = df.select("start_time","end_time","stats.*")

内容的提问来源于stack exchange,提问作者wavemonger

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 10:45:04