You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

拆分行级JSON超大文件后结构异常、出现转义字符问题求助

问题根因

你遇到的两个异常都是json.dump的错误使用导致的:

  • 外层嵌套数组:你将逐行解析得到的单个数组对象存入了一个大列表,再对大列表整体做序列化,自然会生成外层嵌套的数组结构
  • 多余引号与转义符:你直接把每行已经是合法JSON格式的字符串当成普通值传入json.dump,JSON序列化机制会自动给字符串加双引号、转义内部的特殊字符,相当于执行了两次序列化操作
最优解决方案(无JSON解析,性能最高,适合超大文件)

不需要对JSON内容做任何解析,直接逐行读写即可,100%保留原文件格式,内存占用极低:

# 拆分配置
per_file_lines = 1000  # 每个拆分文件的行数
input_path = "large_file.jsonl"  # 你的原始大文件路径
output_pattern = "split_part_{}.jsonl"  # 拆分文件命名规则

current_line_count = 0
output_file = None

try:
    with open(input_path, "r", encoding="utf-8") as input_f:
        for raw_line in input_f:
            # 达到行数阈值时切换新的输出文件
            if current_line_count % per_file_lines == 0:
                if output_file is not None:
                    output_file.close()
                part_index = current_line_count // per_file_lines + 1
                output_file = open(output_pattern.format(part_index), "w", encoding="utf-8")
            # 直接写入原行内容,不做任何序列化处理
            output_file.write(raw_line)
            current_line_count += 1
finally:
    # 确保最后一个文件正常关闭
    if output_file is not None:
        output_file.close()
可选方案(需校验JSON合法性场景)

如果你需要在拆分时验证每行JSON的有效性,可以用如下实现:

import json

per_file_lines = 1000
input_path = "large_file.jsonl"
output_pattern = "split_part_{}.jsonl"

current_line_count = 0
output_file = None

try:
    with open(input_path, "r", encoding="utf-8") as input_f:
        for raw_line in input_f:
            raw_line = raw_line.strip()
            if not raw_line:
                continue
            # 校验JSON格式合法性
            json_obj = json.loads(raw_line)
            if current_line_count % per_file_lines == 0:
                if output_file is not None:
                    output_file.close()
                part_index = current_line_count // per_file_lines + 1
                output_file = open(output_pattern.format(part_index), "w", encoding="utf-8")
            # 单条序列化后加换行写入,和原格式对齐
            output_file.write(json.dumps(json_obj, ensure_ascii=False) + "\n")
            current_line_count += 1
finally:
    if output_file is not None:
        output_file.close()

注意:不要把多行解析后的JSON对象存入列表后整体序列化,必须逐行单独序列化写入才能保留原每行一个JSON数组的结构。

内容的提问来源于stack exchange,提问作者Avv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 03:24:03