You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark中调用query_json后RDD为何返回空数组?

问题原因与修复方案

核心问题:路径匹配逻辑不匹配

你在PySpark中得到空数组的根本原因是传入的JSON片段和路径前缀不匹配:

  • 调用query_json(x.body, path)时,传入的json_data已经是原JSON里root.body对应的内容
  • 但find函数默认的初始路径是'root',遍历json_data的子节点时,生成的路径是root.xxx,而你的目标路径是root.body.data.children.data.body,两者完全无法匹配,自然收集不到任何结果
  • 本地Python内核能正常运行,是因为你传入的是完整的JSON对象(包含root层级),路径前缀能对应上

修复方法(二选一即可)

方法一:调整初始路径前缀

修改query_json函数中调用find的语句,让初始路径和传入的JSON片段对应:

def query_json(json_data, path):
    nodes = []

    def find(obj, current_path='root'):
        if current_path == path:
            nodes.append(obj)
        elif isinstance(obj, dict):
            for key, value in obj.items():
                find(value, f'{current_path}.{key}')
        elif isinstance(obj, list):
            for index, value in enumerate(obj):
                find(value, f'{current_path}')

    # 调整初始current_path为root.body,匹配传入的x.body层级
    find(json_data, current_path='root.body')
    return nodes

path = 'root.body.data.children.data.body'
df = json_df.limit(1)
rdd = df.rdd.map(lambda x: query_json(x.body, path))
rdd.collect()

方法二:修改目标路径

既然传入的是x.body(即原root.body部分),直接把目标路径的root.body前缀去掉:

path = 'root.data.children.data.body'
# 其余代码保持不变

额外提示

处理数组时,你的代码没有给路径添加索引(比如root.body.data.children[0]),不过因为目标路径里也没有索引,所以只要数组元素的子结构符合data.body,就能正常收集所有匹配的内容,这部分逻辑是没问题的。

内容的提问来源于stack exchange,提问作者matt-in-the-hat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 20:53:08