拆分行级JSON超大文件后结构异常、出现转义字符问题求助
问题根因
你遇到的两个异常都是json.dump的错误使用导致的:
- 外层嵌套数组:你将逐行解析得到的单个数组对象存入了一个大列表,再对大列表整体做序列化,自然会生成外层嵌套的数组结构
- 多余引号与转义符:你直接把每行已经是合法JSON格式的字符串当成普通值传入
json.dump,JSON序列化机制会自动给字符串加双引号、转义内部的特殊字符,相当于执行了两次序列化操作
最优解决方案(无JSON解析,性能最高,适合超大文件)
不需要对JSON内容做任何解析,直接逐行读写即可,100%保留原文件格式,内存占用极低:
# 拆分配置 per_file_lines = 1000 # 每个拆分文件的行数 input_path = "large_file.jsonl" # 你的原始大文件路径 output_pattern = "split_part_{}.jsonl" # 拆分文件命名规则 current_line_count = 0 output_file = None try: with open(input_path, "r", encoding="utf-8") as input_f: for raw_line in input_f: # 达到行数阈值时切换新的输出文件 if current_line_count % per_file_lines == 0: if output_file is not None: output_file.close() part_index = current_line_count // per_file_lines + 1 output_file = open(output_pattern.format(part_index), "w", encoding="utf-8") # 直接写入原行内容,不做任何序列化处理 output_file.write(raw_line) current_line_count += 1 finally: # 确保最后一个文件正常关闭 if output_file is not None: output_file.close()
可选方案(需校验JSON合法性场景)
如果你需要在拆分时验证每行JSON的有效性,可以用如下实现:
import json per_file_lines = 1000 input_path = "large_file.jsonl" output_pattern = "split_part_{}.jsonl" current_line_count = 0 output_file = None try: with open(input_path, "r", encoding="utf-8") as input_f: for raw_line in input_f: raw_line = raw_line.strip() if not raw_line: continue # 校验JSON格式合法性 json_obj = json.loads(raw_line) if current_line_count % per_file_lines == 0: if output_file is not None: output_file.close() part_index = current_line_count // per_file_lines + 1 output_file = open(output_pattern.format(part_index), "w", encoding="utf-8") # 单条序列化后加换行写入,和原格式对齐 output_file.write(json.dumps(json_obj, ensure_ascii=False) + "\n") current_line_count += 1 finally: if output_file is not None: output_file.close()
注意:不要把多行解析后的JSON对象存入列表后整体序列化,必须逐行单独序列化写入才能保留原每行一个JSON数组的结构。
内容的提问来源于stack exchange,提问作者Avv
相关产品推荐
相关产品推荐

