PySpark中调用query_json后RDD为何返回空数组?
问题原因与修复方案
核心问题:路径匹配逻辑不匹配
你在PySpark中得到空数组的根本原因是传入的JSON片段和路径前缀不匹配:
- 调用
query_json(x.body, path)时,传入的json_data已经是原JSON里root.body对应的内容 - 但
find函数默认的初始路径是'root',遍历json_data的子节点时,生成的路径是root.xxx,而你的目标路径是root.body.data.children.data.body,两者完全无法匹配,自然收集不到任何结果 - 本地Python内核能正常运行,是因为你传入的是完整的JSON对象(包含
root层级),路径前缀能对应上
修复方法(二选一即可)
方法一:调整初始路径前缀
修改query_json函数中调用find的语句,让初始路径和传入的JSON片段对应:
def query_json(json_data, path): nodes = [] def find(obj, current_path='root'): if current_path == path: nodes.append(obj) elif isinstance(obj, dict): for key, value in obj.items(): find(value, f'{current_path}.{key}') elif isinstance(obj, list): for index, value in enumerate(obj): find(value, f'{current_path}') # 调整初始current_path为root.body,匹配传入的x.body层级 find(json_data, current_path='root.body') return nodes path = 'root.body.data.children.data.body' df = json_df.limit(1) rdd = df.rdd.map(lambda x: query_json(x.body, path)) rdd.collect()
方法二:修改目标路径
既然传入的是x.body(即原root.body部分),直接把目标路径的root.body前缀去掉:
path = 'root.data.children.data.body' # 其余代码保持不变
额外提示
处理数组时,你的代码没有给路径添加索引(比如root.body.data.children[0]),不过因为目标路径里也没有索引,所以只要数组元素的子结构符合data.body,就能正常收集所有匹配的内容,这部分逻辑是没问题的。
内容的提问来源于stack exchange,提问作者matt-in-the-hat
相关产品推荐
相关产品推荐

