在Python中流式读取超大JSON.zstd压缩文件的方案
超大.json.zstd文件分块迭代读取实现方案
核心逻辑
基于zstandard库的流式解压能力,逐块读取压缩文件并实时解压,同时按可配置的块大小拆分解压后的JSON内容,确保无需加载整个文件到内存。针对不同的写入格式(行分隔JSON单元/连续JSON数组),提供两种实现方式。
实现代码
场景1:写入时按行存储JSON单元(如每个data_chunk是一行JSON)
import zstandard as zstd def iter_zstd_line_chunks(file_path, chunk_size=1024*1024): with open(file_path, 'rb') as f: dctx = zstd.ZstdDecompressor() with dctx.stream_reader(f) as reader: remaining = b'' while True: chunk = reader.read(chunk_size) if not chunk: # 输出剩余的最后一行 if remaining: yield remaining.decode('utf-8') break # 合并剩余内容与当前块,按换行拆分 full_chunk = remaining + chunk lines = full_chunk.split(b'\n') # 输出除最后一行外的所有完整行 for line in lines[:-1]: if line: # 跳过空行 yield line.decode('utf-8') # 保留最后一行,用于下一次拼接 remaining = lines[-1]
场景2:写入时为连续JSON数组(如[obj1, obj2, obj3,...])
import zstandard as zstd import json def iter_zstd_array_chunks(file_path, chunk_size=1024*1024): with open(file_path, 'rb') as f: dctx = zstd.ZstdDecompressor() with dctx.stream_reader(f) as reader: buffer = b'' open_braces = 0 skip_start_bracket = True while True: chunk = reader.read(chunk_size) if not chunk: # 输出剩余的完整JSON对象 if buffer.strip(): yield buffer.decode('utf-8').strip() break buffer += chunk # 跳过开头的数组左括号 if skip_start_bracket: if b'[' in buffer: buffer = buffer[buffer.index(b'[')+1:] skip_start_bracket = False else: continue # 遍历buffer,按完整JSON对象拆分 idx = 0 while idx < len(buffer): if buffer[idx] == ord(b'{'): open_braces += 1 elif buffer[idx] == ord(b'}'): open_braces -= 1 # 找到闭合的JSON对象,检查后续分隔符 if open_braces == 0: end_idx = idx + 1 # 跳过逗号或数组右括号 while end_idx < len(buffer) and buffer[end_idx] in (ord(b','), ord(b']')): end_idx += 1 # 输出完整对象 yield buffer[:end_idx].decode('utf-8').strip() # 截断buffer,处理剩余内容 buffer = buffer[end_idx:] idx = 0 continue idx += 1
使用示例
# 处理行分隔JSON单元 for line in iter_zstd_line_chunks('large_data.json.zstd', chunk_size=4*1024*1024): data = json.loads(line) # 业务逻辑处理 print(data['id']) # 处理JSON数组格式 for obj_str in iter_zstd_array_chunks('large_data.json.zstd', chunk_size=8*1024*1024): data = json.loads(obj_str) # 业务逻辑处理 print(data['name'])
关键说明
chunk_size参数可根据内存资源灵活调整,数值越大,单次处理的数据量越多,内存占用越高- 流式解压通过
zstd.stream_reader实现,仅在内存中保留当前处理的块数据,无需加载整个压缩文件 - 两种场景分别对应不同的写入格式,可根据实际写入的
data_chunks结构选择合适的实现
内容的提问来源于stack exchange,提问作者jeandut
相关产品推荐
相关产品推荐

