Python如何将文件中多个逗号分隔的JSON对象合并为列表
Python 转换逗号分隔JSON对象为JSON数组方法
你提到的存储格式本质是缺少首尾中括号的JSON数组,根据文件大小和语法规范程度可以选择以下处理方案:
方案1:直接拼接中括号解析(适合小文件、内容无语法异常)
仅需要Python内置的json模块即可完成,不需要额外安装依赖:
import json # 读取原文件内容 with open("input.json", "r", encoding="utf-8") as f: raw_content = f.read() # 替换示例中可能存在的中文全角引号(如果你的文件已经是英文引号可删除此行) raw_content = raw_content.replace('“', '"').replace('”', '"') # 首尾拼接中括号转成标准JSON数组格式 json_list = json.loads(f"[{raw_content}]") # 如需导出为标准JSON文件可执行以下代码 with open("output.json", "w", encoding="utf-8") as f: json.dump(json_list, f, indent=2, ensure_ascii=False)
方案2:兼容宽松语法解析(适合存在多余逗号、格式不严谨的场景)
如果文件末尾存在多余逗号、或者有其他非标准JSON语法,可以使用json5库解析,它支持宽松的JSON语法规则:
- 先安装依赖:
pip install json5
- 处理代码:
import json import json5 with open("input.json", "r", encoding="utf-8") as f: raw_content = f.read() raw_content = raw_content.replace('“', '"').replace('”', '"') json_list = json5.loads(f"[{raw_content}]") # 导出标准JSON with open("output.json", "w", encoding="utf-8") as f: json.dump(json_list, f, indent=2, ensure_ascii=False)
方案3:大文件增量解析(适合GB级超大文件,避免内存溢出)
如果文件体积过大无法一次性读入内存,可以使用ijson库做增量解析,逐对象读取后存入列表:
- 安装依赖:
pip install ijson
- 处理代码:
import json import ijson # 先临时生成带中括号的临时文件 with open("input.json", "r", encoding="utf-8") as in_f, open("temp.json", "w", encoding="utf-8") as out_f: out_f.write("[") for line in in_f: line = line.replace('“', '"').replace('”', '"') out_f.write(line) out_f.write("]") # 增量读取所有对象 json_list = [] with open("temp.json", "r", encoding="utf-8") as f: for item in ijson.items(f, "item"): json_list.append(item) # 导出结果 with open("output.json", "w", encoding="utf-8") as f: json.dump(json_list, f, indent=2, ensure_ascii=False)
内容的提问来源于stack exchange,提问作者rashxxx
相关产品推荐
相关产品推荐

