You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark DataFrame中遍历非空嵌套JSON字段并提取指定字段?

处理PySpark DataFrame中嵌套JSON字段的非空情况

针对你的需求,我们可以在原有代码基础上添加非空children字段的处理逻辑,遍历每个嵌套子结构并提取数据:

data = []
columns = ['created_date','description','id','last_modified_date','links','name','order','parent_id','pid','recursive','project_id']

for row in landing_df.collect():
    # 处理children为空数组的情况
    if len(row["children"]) == 0:
        data.append([
            row["created_date"],
            row["description"],
            row["id"],
            row["last_modified_date"],
            row["links"],
            row["name"],
            row["order"],
            -1,
            row["pid"],
            row["recursive"],
            row["project_id"]
        ])
        print(row["id"])
    # 处理children非空的情况
    else:
        # 遍历每个子结构
        for child in row["children"]:
            # 提取子结构对应字段,parent_id设为当前行的id
            data.append([
                child.get("created_date"),  # 用get避免字段缺失报错
                child.get("description"),
                child.get("id"),
                child.get("last_modified_date"),
                child.get("links"),
                child.get("name"),
                child.get("order"),
                row["id"],  # 父行id作为子结构的parent_id
                child.get("pid"),
                child.get("recursive"),
                child.get("project_id")
            ])

sourceDF1 = spark.createDataFrame(data, columns)

关键说明:

  • 保留原有空数组处理逻辑,新增非空场景的遍历逻辑
  • 使用child.get()替代直接索引,避免子结构字段缺失导致报错
  • 非空场景下将父行id设为子结构的parent_id,与空场景的-1形成区分

调整提示:

如果children子结构的字段名和父行不完全匹配,直接修改child.get()中的字段名即可;若不需要保留父行其他字段,可按需调整append的列表内容。

内容的提问来源于stack exchange,提问作者tathagat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 04:15:27