如何在PySpark DataFrame中遍历非空嵌套JSON字段并提取指定字段?
处理PySpark DataFrame中嵌套JSON字段的非空情况
针对你的需求,我们可以在原有代码基础上添加非空children字段的处理逻辑,遍历每个嵌套子结构并提取数据:
data = [] columns = ['created_date','description','id','last_modified_date','links','name','order','parent_id','pid','recursive','project_id'] for row in landing_df.collect(): # 处理children为空数组的情况 if len(row["children"]) == 0: data.append([ row["created_date"], row["description"], row["id"], row["last_modified_date"], row["links"], row["name"], row["order"], -1, row["pid"], row["recursive"], row["project_id"] ]) print(row["id"]) # 处理children非空的情况 else: # 遍历每个子结构 for child in row["children"]: # 提取子结构对应字段,parent_id设为当前行的id data.append([ child.get("created_date"), # 用get避免字段缺失报错 child.get("description"), child.get("id"), child.get("last_modified_date"), child.get("links"), child.get("name"), child.get("order"), row["id"], # 父行id作为子结构的parent_id child.get("pid"), child.get("recursive"), child.get("project_id") ]) sourceDF1 = spark.createDataFrame(data, columns)
关键说明:
- 保留原有空数组处理逻辑,新增非空场景的遍历逻辑
- 使用
child.get()替代直接索引,避免子结构字段缺失导致报错 - 非空场景下将父行
id设为子结构的parent_id,与空场景的-1形成区分
调整提示:
如果children子结构的字段名和父行不完全匹配,直接修改child.get()中的字段名即可;若不需要保留父行其他字段,可按需调整append的列表内容。
内容的提问来源于stack exchange,提问作者tathagat
相关产品推荐
相关产品推荐

