Python循环优化:大体积NDJSON文件处理提速方案咨询
优化NDJSON大文件处理速度的方案
核心瓶颈分析
你的代码当前速度慢的主要原因有三个:
- 循环内反复打开/关闭输出文件:每次文件操作都会触发系统调用,IO开销极大
- 单线程执行:没有利用多核CPU资源,导致CPU使用率仅8.5%
- 内存未充分利用:64GB内存只用到3.5GB,说明数据加载和处理模式没有最大化内存优势
具体优化步骤
1. 彻底修复文件IO的低效问题
把输出文件的打开操作移到循环外部,避免反复创建文件句柄:
# 提前打开输出文件,循环内仅执行写入操作 with open('file_name.ndjson', 'a') as out_file: for json_obj in json_list: some_scipy_function() if condition_is_satisfied: json_obj['new_key'] = new_value out_file.write(json.dumps(json_obj) + '\n') # 可选:若担心缓冲区积压,可定期刷新 # out_file.flush()
注意:如果是生成全新输出文件,建议用'w'模式替代'a',避免重复追加历史内容
2. 利用多核CPU并行处理
some_scipy_function属于CPU密集型任务,Python的GIL会限制单线程性能,因此用多进程来利用多核:
方案A:全量加载+进程池批量处理
如果64GB内存足够容纳3.5GB的NDJSON数据,直接全量加载后并行处理:
from multiprocessing import Pool import json def process_single_obj(json_obj): some_scipy_function() if condition_is_satisfied: json_obj['new_key'] = new_value return json.dumps(json_obj) + '\n' return None # 一次性加载所有数据到内存 with open('input.ndjson', 'r') as in_file: json_list = [json.loads(line) for line in in_file] # 进程数设为CPU核心数(比如8核就填8) with Pool(processes=8) as pool: results = pool.map(process_single_obj, json_list) # 过滤无效结果,批量写入文件 with open('file_name.ndjson', 'w') as out_file: out_file.writelines(filter(None, results))
方案B:分块加载+并行处理(内存友好型)
如果担心全量加载内存压力过大,可分批次处理:
from concurrent.futures import ProcessPoolExecutor import json def process_obj(obj): some_scipy_function() if condition_is_satisfied: obj['new_key'] = new_value return json.dumps(obj) + '\n' return None # 每批次处理10万条数据,可根据内存调整 batch_size = 100000 with open('input.ndjson', 'r') as in_file, open('file_name.ndjson', 'w') as out_file: batch = [] for line in in_file: batch.append(json.loads(line)) if len(batch) >= batch_size: with ProcessPoolExecutor(max_workers=8) as executor: for result in executor.map(process_obj, batch): if result: out_file.write(result) batch = [] # 处理剩余的小批次数据 if batch: with ProcessPoolExecutor(max_workers=8) as executor: for result in executor.map(process_obj, batch): if result: out_file.write(result)
3. 最大化内存利用,减少IO次数
- 尽可能一次性加载输入文件到内存:64GB内存完全可以容纳3.5GB的NDJSON,避免反复读取磁盘
- 批量收集符合条件的结果,再一次性写入:减少文件写入的系统调用次数
4. 细节优化进一步提速
- 替换JSON库:用
ujson或orjson替代标准库json,序列化/反序列化速度提升2-5倍import ujson # 需先执行 pip install ujson # 替换json.dumps为ujson.dumps,json.loads为ujson.loads - 向量化处理:如果
some_scipy_function支持向量化,可批量提取所有json_obj的相关数据,用scipy向量化函数一次性处理,比单条处理效率高一个数量级 - 简化条件判断:尽量减少循环内的计算逻辑,把非必要的计算移到循环外部
内容的提问来源于stack exchange,提问作者ennezetaqu
相关产品推荐
相关产品推荐

