You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark Structured Streaming:将JSON结构体数组拆分为DataFrame行

解决方案

你需要使用Spark的explode函数将结构体数组拆分为独立行,修改后的代码如下:

df.select(from_json($"col", schemaAsJson) as "json")
  .select(explode($"json") as "customer_info") // 将数组中的每个结构体拆分为单独行
  .select("customer_info.customer", "customer_info.sex", "customer_info.country")

代码说明:

  1. 解析JSON后得到的json列是结构体数组类型,explode($"json")会将数组中的每个结构体元素单独生成一行记录,并重命名为customer_info。
  2. 从炸开后的customer_info结构体中提取字段,即可得到每行一条用户数据的结果。

执行这段代码后,输出结果会和你期望的一致:

+--------------+----------------+----------------+
|      customer|             sex|         country|
+--------------+----------------+----------------+
|           Jim|            male|              US|
|           Pam|          female|              US|
+--------------+----------------+----------------+

内容的提问来源于stack exchange,提问作者Harikrishnan Balachandran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 10:40:29