如何用Python遍历大型JSON生成精简输出及遍历含子节点的JSON
问题1:遍历大型JSON文件并生成精简JSON输出
处理大型JSON文件的核心是避免一次性加载整个文件到内存,推荐用ijson库做流式解析,只提取需要的字段后写入新文件。
步骤:
- 安装依赖库:
pip install ijson
- 流式解析并提取指定字段:
假设你的大型JSON是数组结构(如[{"id":1, "name":"a", "other":"xxx"}, ...]),要提取id和name字段:
import ijson import json input_file = "large_data.json" output_file = "trimmed_data.json" with open(input_file, "rb") as f_in, open(output_file, "w") as f_out: # 写入JSON数组开头 f_out.write("[") first_item = True # 流式遍历每个元素 for item in ijson.items(f_in, "item"): # 构建精简对象 trimmed_item = { "id": item.get("id"), "name": item.get("name") } # 处理元素间的逗号分隔 if not first_item: f_out.write(",") else: first_item = False # 写入当前精简对象 json.dump(trimmed_item, f_out) # 写入JSON数组结尾 f_out.write("]")
说明:
ijson.items(f_in, "item")会逐个解析JSON数组中的元素,不会加载整个文件,适合GB级别的JSON文件。- 若你的JSON是单个嵌套对象而非数组,调整
ijson.items的路径参数即可(如""表示整个对象)。
问题2:遍历带有子节点的JSON数据
嵌套JSON的遍历通常用递归函数处理,不管是已加载到内存的JSON对象,还是流式解析的节点都适用。
场景1:遍历已加载的嵌套JSON(小文件)
如果JSON文件不大,直接用json.load()加载后递归遍历:
import json def traverse_json(node, path=""): if isinstance(node, dict): for key, value in node.items(): current_path = f"{path}.{key}" if path else key if isinstance(value, (dict, list)): traverse_json(value, current_path) else: print(f"{current_path}: {value}") elif isinstance(node, list): for index, item in enumerate(node): current_path = f"{path}[{index}]" traverse_json(item, current_path) # 加载JSON数据 with open("nested_data.json", "r") as f: data = json.load(f) # 开始遍历 traverse_json(data)
场景2:流式遍历嵌套JSON(大文件)
如果是大型嵌套JSON,结合ijson的路径匹配递归解析子节点:
import ijson def traverse_large_nested_json(file_path): with open(file_path, "rb") as f: # 遍历所有层级的键值对 for prefix, event, value in ijson.parse(f): if event in ("string", "number", "boolean", "null"): print(f"{prefix}: {value}") traverse_large_nested_json("large_nested_data.json")
说明:
- 递归函数会自动处理
dict(对象)和list(数组)类型的子节点,遍历所有层级的键值对。 - 流式解析时,
ijson.parse()会返回每个节点的路径、事件类型和值,无需加载整个文件。
内容的提问来源于stack exchange,提问作者paliknight
相关产品推荐
相关产品推荐

