You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分20GB大型JSON文件为100MB小文件?选Shell还是Python?

拆分大型JSON数组文件的最佳方案:Python vs Shell

结论:优先选择Python实现

Shell命令无法妥善处理JSON的结构化语法,容易生成无效的JSON文件;Python可以流式解析超大文件,精准控制拆分后的文件大小,同时保证每个输出文件都是合法的JSON数组。

为什么不推荐Shell命令?

常用的Shell拆分工具(如split)是基于字节或行数拆分的,而JSON数组的元素是结构化数据——拆分时很可能把一个完整的JSON对象拆成两半,导致输出文件语法错误。即使手动处理每个文件的首尾括号,也无法精准控制每个文件的元素数量和最终大小,操作繁琐且不可靠。

Python实现方案(低内存、高可靠)

针对20GB级别的超大JSON数组,必须采用流式解析避免内存溢出,这里推荐使用ijson库(轻量级流式JSON解析器)。

步骤1:安装依赖

pip install ijson

步骤2:拆分脚本

import os
import ijson

def split_large_json(input_file, max_size_mb=100):
    max_size = max_size_mb * 1024 * 1024  # 转换为字节单位
    file_index = 1
    current_file = None
    current_size = 0
    is_first_elem = True

    with open(input_file, 'rb') as src_file:
        # 流式读取原JSON数组中的每个元素
        elements = ijson.items(src_file, 'item')
        for elem in elements:
            # 将元素转为合法的JSON字符串(替换单引号为双引号)
            elem_str = f"{',' if not is_first_elem else ''}{repr(elem).replace(\"'\", '\"')}"
            elem_bytes_size = len(elem_str.encode('utf-8'))

            # 初始化新文件
            if not current_file:
                output_name = f"{os.path.splitext(input_file)[0]}{file_index}.json"
                current_file = open(output_name, 'w', encoding='utf-8')
                current_file.write('[')
                current_size = 1  # 记录开头'['的字节数
                is_first_elem = True

            # 检查当前文件添加元素后是否超过大小限制
            if current_size + elem_bytes_size > max_size:
                # 关闭当前文件,补全JSON数组结尾
                current_file.write(']')
                current_file.close()
                # 新建下一个文件
                file_index += 1
                output_name = f"{os.path.splitext(input_file)[0]}{file_index}.json"
                current_file = open(output_name, 'w', encoding='utf-8')
                current_file.write('[')
                current_size = 1
                is_first_elem = True
                elem_str = repr(elem).replace("'", '"')  # 新文件第一个元素不需要前置逗号

            # 写入当前元素
            current_file.write(elem_str)
            current_size += elem_bytes_size
            is_first_elem = False

        # 处理最后一个文件的结尾
        if current_file:
            current_file.write(']')
            current_file.close()

if __name__ == '__main__':
    split_large_json('file.json', max_size_mb=100)

脚本说明

  • 流式解析:通过ijson.items逐个读取数组元素,内存占用始终保持在低水平,不会加载整个20GB文件到内存。
  • 大小控制:实时计算当前文件的字节大小,接近100MB时自动切换到新文件。
  • 语法保证:每个输出文件都会自动添加[开头和]结尾,元素间用逗号分隔,确保是合法的JSON数组。

无第三方库替代方案(标准库实现)

如果无法安装第三方库,可以使用json.JSONDecoder手动流式解析,但代码会更繁琐,核心思路仍是逐元素读取并控制文件大小。


内容的提问来源于stack exchange,提问作者kites

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 15:35:12