Python追加多字典到JSON文件格式错误的解决方案咨询
Python低内存占用JSON多字典追加写入方案
问题根源
直接循环调用json.dump以追加模式写入独立JSON对象,会导致多个顶层对象拼接时出现}{无逗号分隔的问题,完全不符合JSON语法规范。
实现方案
因业务数据量过大无法全量加载到内存组装为单个字典,采用手动维护JSON顶层结构+单对象流式写入的方案,内存仅占用单条repo数据的大小,满足性能要求:
实现步骤
- 初始化阶段:先向空文件写入JSON顶层左大括号
{ - 写入阶段:新增标记位判断是否为首次写入,首次写入repo数据直接序列化写入,非首次写入前先追加逗号
,做分隔 - 收尾阶段:所有数据写入完成后,单独写入顶层右大括号
}完成JSON闭合
修正后代码
import json from tqdm import tqdm # 初始化JSON文件,写入顶层左大括号 with open("repoPrDetails.json", mode="w", encoding="utf-8") as file: file.write("{") first_write = True # 首次写入标记位 for repoName in tqdm(repoList, total=len(repoList), desc="extracting PR details"): consolidated_JSON = {} end_cursor, has_next_page = None, True repoPR = {} while has_next_page: data = getPrJSON(repoName, end_cursor) # 此处补充原有分页逻辑,比如更新end_cursor、has_next_page状态的代码 for pr in data["nodes"]: repoPrDetails = { "PR_number": pr["number"], "title": pr["title"], "id": pr["id"], } consolidated_JSON[f"PR{pr['number']}"] = repoPrDetails repoPR[repoName] = consolidated_JSON # 流式写入当前repo数据 with open("repoPrDetails.json", mode="a", encoding="utf-8") as file: if not first_write: file.write(",") # 非首次写入先加逗号分隔 json.dump(repoPR, file, default=str, ensure_ascii=False) first_write = False # 写入顶层右大括号完成JSON闭合 with open("repoPrDetails.json", mode="a", encoding="utf-8") as file: file.write("}")
最终合法JSON结构示例
{ "repo1": { "PR301": { "PR_number": 301, "title": "Update configuration.json", "id": "MDExOlB1bGxSZXF1ZXN0NDQ1MTM1" }, "PR302": { "PR_number": 302, "title": "refactor", "id": "MDExOlB1bGxSZXF1ZXN0NDQ1MTM2" } }, "repo2": { "PR301": { "PR_number": 301, "title": "Update configuration.json", "id": "MDExOlB1bGxSZXF1ZXN0NDQ1MTM1" } } }
内容的提问来源于stack exchange,提问作者Aryan Saxena
相关产品推荐
相关产品推荐

