Spark Structured Streaming:将JSON结构体数组拆分为DataFrame行
解决方案
你需要使用Spark的explode函数将结构体数组拆分为独立行,修改后的代码如下:
df.select(from_json($"col", schemaAsJson) as "json") .select(explode($"json") as "customer_info") // 将数组中的每个结构体拆分为单独行 .select("customer_info.customer", "customer_info.sex", "customer_info.country")
代码说明:
- 解析JSON后得到的
json列是结构体数组类型,explode($"json")会将数组中的每个结构体元素单独生成一行记录,并重命名为customer_info。 - 从炸开后的
customer_info结构体中提取字段,即可得到每行一条用户数据的结果。
执行这段代码后,输出结果会和你期望的一致:
+--------------+----------------+----------------+ | customer| sex| country| +--------------+----------------+----------------+ | Jim| male| US| | Pam| female| US| +--------------+----------------+----------------+
内容的提问来源于stack exchange,提问作者Harikrishnan Balachandran
相关产品推荐
相关产品推荐

