You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多线程池JSON批次低内存流式写入单文件方案咨询

安全流式写入批量JSON数据到文件的方案

问题背景

使用Python的ThreadPool多线程池处理任务时,每个任务会返回一批JSON对象数组,总数据量约150万条、1.5GB,无法将完整结果存入内存,需要写入单个合法的JSON文件。试过json-stream库但不支持批次嵌套的列表结构,也考虑过手动写入[、分隔逗号、]的字符串操作,希望找到更安全的库实现方式。

推荐安全方案

方案1:结合jsonlines构建合法JSON数组

jsonlines库专门用于逐行处理JSON对象,我们可以基于它来构建完整的JSON数组,既保证序列化安全,又避免内存溢出:

from multiprocessing.pool import ThreadPool
import jsonlines
from tqdm import tqdm

def thread(worker, jobs, threads):
    pool = ThreadPool(threads)
    is_first_element = True
    with open("tmp.json", "w") as f, jsonlines.Writer(f) as writer:
        # 写入JSON数组开头
        f.write("[")
        for result in tqdm(pool.imap_unordered(worker, jobs), total=len(jobs)):
            for element in result:
                if not is_first_element:
                    # 非首个元素,先写分隔逗号+换行
                    f.write(",\n")
                else:
                    is_first_element = False
                # 用jsonlines安全序列化单个对象并写入
                writer.write(element)
        # 写入JSON数组结尾
        f.write("]")

方案2:用json.dump优化手动流式写入

如果不想额外安装库,用标准库的json.dump代替手动拼接JSON字符串,既保留流式写入的内存优势,又避免手动序列化的语法错误:

from multiprocessing.pool import ThreadPool
import json
from tqdm import tqdm

def thread(worker, jobs, threads):
    pool = ThreadPool(threads)
    is_first = True
    with open("tmp.json", "w") as f:
        f.write("[")
        for result in tqdm(pool.imap_unordered(worker, jobs), total=len(jobs)):
            for elem in result:
                if not is_first:
                    f.write(", ")
                # 用标准库json.dump安全序列化单个对象
                json.dump(elem, f)
                is_first = False
        f.write("]")

方案优势

  • 避免手动拼接JSON字符串可能出现的转义错误、逗号位置错误等问题
  • 利用成熟库/标准库的JSON序列化逻辑,保证每个对象的格式合法性
  • 全程流式写入,仅在内存中保留单条JSON对象,完全适配大内存压力场景

内容的提问来源于stack exchange,提问作者OJT

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 14:34:59