You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Python中流式读取超大JSON.zstd压缩文件的方案

超大.json.zstd文件分块迭代读取实现方案

核心逻辑

基于zstandard库的流式解压能力,逐块读取压缩文件并实时解压,同时按可配置的块大小拆分解压后的JSON内容,确保无需加载整个文件到内存。针对不同的写入格式(行分隔JSON单元/连续JSON数组),提供两种实现方式。

实现代码

场景1:写入时按行存储JSON单元(如每个data_chunk是一行JSON)

import zstandard as zstd

def iter_zstd_line_chunks(file_path, chunk_size=1024*1024):
    with open(file_path, 'rb') as f:
        dctx = zstd.ZstdDecompressor()
        with dctx.stream_reader(f) as reader:
            remaining = b''
            while True:
                chunk = reader.read(chunk_size)
                if not chunk:
                    # 输出剩余的最后一行
                    if remaining:
                        yield remaining.decode('utf-8')
                    break
                # 合并剩余内容与当前块,按换行拆分
                full_chunk = remaining + chunk
                lines = full_chunk.split(b'\n')
                # 输出除最后一行外的所有完整行
                for line in lines[:-1]:
                    if line:  # 跳过空行
                        yield line.decode('utf-8')
                # 保留最后一行,用于下一次拼接
                remaining = lines[-1]

场景2:写入时为连续JSON数组(如[obj1, obj2, obj3,...])

import zstandard as zstd
import json

def iter_zstd_array_chunks(file_path, chunk_size=1024*1024):
    with open(file_path, 'rb') as f:
        dctx = zstd.ZstdDecompressor()
        with dctx.stream_reader(f) as reader:
            buffer = b''
            open_braces = 0
            skip_start_bracket = True
            
            while True:
                chunk = reader.read(chunk_size)
                if not chunk:
                    # 输出剩余的完整JSON对象
                    if buffer.strip():
                        yield buffer.decode('utf-8').strip()
                    break
                
                buffer += chunk
                
                # 跳过开头的数组左括号
                if skip_start_bracket:
                    if b'[' in buffer:
                        buffer = buffer[buffer.index(b'[')+1:]
                        skip_start_bracket = False
                    else:
                        continue
                
                # 遍历buffer,按完整JSON对象拆分
                idx = 0
                while idx < len(buffer):
                    if buffer[idx] == ord(b'{'):
                        open_braces += 1
                    elif buffer[idx] == ord(b'}'):
                        open_braces -= 1
                        # 找到闭合的JSON对象,检查后续分隔符
                        if open_braces == 0:
                            end_idx = idx + 1
                            # 跳过逗号或数组右括号
                            while end_idx < len(buffer) and buffer[end_idx] in (ord(b','), ord(b']')):
                                end_idx += 1
                            # 输出完整对象
                            yield buffer[:end_idx].decode('utf-8').strip()
                            # 截断buffer,处理剩余内容
                            buffer = buffer[end_idx:]
                            idx = 0
                            continue
                    idx += 1

使用示例

# 处理行分隔JSON单元
for line in iter_zstd_line_chunks('large_data.json.zstd', chunk_size=4*1024*1024):
    data = json.loads(line)
    # 业务逻辑处理
    print(data['id'])

# 处理JSON数组格式
for obj_str in iter_zstd_array_chunks('large_data.json.zstd', chunk_size=8*1024*1024):
    data = json.loads(obj_str)
    # 业务逻辑处理
    print(data['name'])

关键说明

  • chunk_size参数可根据内存资源灵活调整,数值越大,单次处理的数据量越多,内存占用越高
  • 流式解压通过zstd.stream_reader实现,仅在内存中保留当前处理的块数据,无需加载整个压缩文件
  • 两种场景分别对应不同的写入格式,可根据实际写入的data_chunks结构选择合适的实现

内容的提问来源于stack exchange,提问作者jeandut

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 15:05:05