如何用Python修复无括号无逗号分隔的无效大JSON文件
修复格式错误的大体积JSON文件(Python实现)
你的JSON文件本质是JSON Lines格式(每行一个独立JSON对象),要转成标准JSON,只需给整体包裹数组括号,并在条目间添加逗号。针对大文件,不能一次性加载到内存,推荐逐行处理:
基础处理代码
直接完成格式转换,适合确定每行都是有效JSON对象的场景:
import json input_file = "待修复的文件.json" output_file = "修复后的文件.json" with open(input_file, 'r', encoding='utf-8') as in_f, open(output_file, 'w', encoding='utf-8') as out_f: out_f.write('[') is_first_entry = True for line in in_f: cleaned_line = line.strip() if not cleaned_line: # 跳过空行 continue if not is_first_entry: out_f.write(',') out_f.write(cleaned_line) is_first_entry = False out_f.write(']')
带格式校验的版本
如果不确定每行是否都是有效JSON,可以加入校验步骤,跳过错误行并提示:
import json input_file = "待修复的文件.json" output_file = "修复后的文件.json" with open(input_file, 'r', encoding='utf-8') as in_f, open(output_file, 'w', encoding='utf-8') as out_f: out_f.write('[') is_first_entry = True for line_num, line in enumerate(in_f, 1): cleaned_line = line.strip() if not cleaned_line: continue try: # 验证当前行是合法JSON json.loads(cleaned_line) except json.JSONDecodeError as err: print(f"第{line_num}行格式错误,已跳过: {err}") continue if not is_first_entry: out_f.write(',') out_f.write(cleaned_line) is_first_entry = False out_f.write(']')
注意事项
- 两种方法都是逐行读写,内存占用极低,适合GB级别的大文件
- 如果你的JSON条目存在跨行情况(比如一个对象占多行),需要先对文件进行预处理合并完整对象,再执行上述转换
内容的提问来源于stack exchange,提问作者José Carlos
相关产品推荐
相关产品推荐

