You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

5GB JSON文件拆分性能优化:提速现有Python处理代码

5GB JSON文件分片优化需求

我有一个5GB的JSON文件,每行是一个独立的JSON对象,需要将其拆分为格式规范、分片大小可配置的多个JSON文件。当前代码因全量读取数据统计记录数,单文件处理耗时约18分钟,急需提速。

输入输出样例

输入样例

{"job": "developer"}
{"job": "taxi driver"}
{"job": "police"}

输出样例

{
  "ROOT": [
    {
      "job": "developer"
    },
    {
      "job": "taxi driver"
    },
    {
      "job": "police"
    }
  ]
}

现有代码

import os
import json
import glob
import time
import shutil
start = time.time()


def filesplit(fInputname,farchivedir,foutdir):
    File_Extension = '.json'
 
    # using partition()
    # String till Substring
    x=foutdir.find(File_Extension)
    res=foutdir[0:x]
    with open(fInputname, 'r', encoding='utf-8') as f1:
        ll = [json.loads(line.strip())  for line in f1.readlines()]

        #Total number of records in the input json file
        print(len(ll))

        #50000 means we getting splits of 50000 json objects
        size_of_the_split=50000
        total = len(ll) // size_of_the_split

        #Number of files getting generated
        print(total+1)
    for i in range(total+1):
        jsonData=ll[i * size_of_the_split:(i + 1) * size_of_the_split]
        json.dump( {'ROOT': jsonData}, open(res + "_" + str(i+1) + ".json", 'w',encoding='utf8'), ensure_ascii=False, indent=True)
    shutil.move(fInputname,farchivedir)
     

for name in glob.glob("C:\\Users\\JSON\\Input\\*.json"):
    print(name)
    filesplit(name, name.replace('C:\\Users\\JSON\\Input','C:\\Users\\JSON\\OriginalFiles_BKP'),name.replace('C:\\Users\\JSON\\Input','C:\\Users\\JSON\\Output'))
  

end = time.time()
print('completed')
print("The time of execution of above program is :",
      (end-start) * 10**3, "ms")

优化方案

现有代码的核心问题是一次性读取全量数据到内存,既占用大量内存,又需要遍历两次(一次读取统计,一次分片写入)。优化思路改为逐行读取、边读边写,无需预先统计总行数,降低内存占用同时减少IO耗时。

优化后代码

import json
import glob
import time
import shutil

def filesplit(f_input, f_archive_dir, f_out_dir, split_size=50000):
    # 生成输出文件前缀
    ext = '.json'
    prefix = f_out_dir[:f_out_dir.find(ext)] if ext in f_out_dir else f_out_dir
    
    file_index = 1
    current_batch = []
    
    with open(f_input, 'r', encoding='utf-8') as infile:
        for line in infile:
            line = line.strip()
            if not line:
                continue  # 跳过空行
            try:
                obj = json.loads(line)
                current_batch.append(obj)
                
                # 当批次达到指定大小,写入文件
                if len(current_batch) == split_size:
                    output_path = f"{prefix}_{file_index}.json"
                    with open(output_path, 'w', encoding='utf-8') as outfile:
                        json.dump({'ROOT': current_batch}, outfile, ensure_ascii=False, indent=True)
                    print(f"已写入文件: {output_path}")
                    current_batch = []
                    file_index += 1
            except json.JSONDecodeError as e:
                print(f"解析JSON行失败: {e}, 行内容: {line}")
                continue
    
    # 处理剩余不足一个批次的数据
    if current_batch:
        output_path = f"{prefix}_{file_index}.json"
        with open(output_path, 'w', encoding='utf-8') as outfile:
            json.dump({'ROOT': current_batch}, outfile, ensure_ascii=False, indent=True)
        print(f"已写入文件: {output_path}")
    
    # 归档原文件
    shutil.move(f_input, f_archive_dir)
    print(f"原文件已归档至: {f_archive_dir}")

if __name__ == "__main__":
    start = time.time()
    # 遍历输入目录下的JSON文件
    for input_path in glob.glob("C:\\Users\\JSON\\Input\\*.json"):
        print(f"开始处理文件: {input_path}")
        archive_path = input_path.replace('C:\\Users\\JSON\\Input', 'C:\\Users\\JSON\\OriginalFiles_BKP')
        output_path = input_path.replace('C:\\Users\\JSON\\Input', 'C:\\Users\\JSON\\Output')
        filesplit(input_path, archive_path, output_path, split_size=50000)
    
    end = time.time()
    print("处理完成")
    print(f"总执行时间: {(end - start) * 10**3:.2f} ms")

优化点说明

  1. 逐行读取处理:避免一次性加载5GB数据到内存,内存占用仅为当前批次的大小(约50000个JSON对象),大幅降低内存压力。
  2. 边读边写:无需预先统计总行数,读取过程中达到批次大小就立即写入文件,减少一次全量遍历的IO耗时。
  3. 错误处理:增加JSON解析错误捕获,避免单个无效行导致整个程序崩溃。
  4. 代码结构优化:使用if __name__ == "__main__"规范入口,参数更清晰,可读性更强。

内容的提问来源于stack exchange,提问作者sk1893

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 06:03:14