如何用Python boto3流式读取S3中的.tar.gz压缩JSON文件?
解决方案:流式处理S3上的大体积.tar.gz JSON文件
一、原生Python实现(无第三方依赖)
直接用boto3+tarfile+json实现流式处理,彻底规避内存溢出和EOF错误:
实现逻辑
- 通过boto3获取S3对象的流式响应体,不下载整个文件到本地
- 用
tarfile的流式模式(r|gz)解压,分块处理压缩数据,不占用大量内存 - 遍历压缩包内的每个文件,逐行读取JSON对象并处理
代码示例
import boto3 import tarfile import json def process_s3_targz_json(bucket_name, s3_key): s3 = boto3.client('s3') # 获取流式响应,避免一次性下载大文件 response = s3.get_object(Bucket=bucket_name, Key=s3_key) stream = response['Body'] # 流式解压tar.gz,r|gz模式支持边读边解 with tarfile.open(fileobj=stream, mode='r|gz') as tar: for member in tar: if member.isfile(): file_obj = tar.extractfile(member) if not file_obj: continue # 逐行读取并解析JSON for line in file_obj: try: # 显式用UTF-8解码,处理非ASCII字符 json_obj = json.loads(line.decode('utf-8').strip()) # 替换为你的业务处理逻辑 # handle_json(json_obj) except json.JSONDecodeError as e: print(f"JSON解析失败: {e}") except UnicodeDecodeError as e: print(f"字符解码失败: {e}") file_obj.close() # 调用示例 process_s3_targz_json('你的存储桶名称', '文件在S3中的路径/xxx.tar.gz')
EOFError解决说明
之前的内存解压方式是一次性把整个压缩文件加载到内存再解压,30GB的解压后文件远超内存容量,导致解压过程中数据截断或内存耗尽,触发EOFError。流式处理是分块读取、分块解压,每块仅占用少量内存,完全避免这个问题。
二、修复smart_open的乱码问题
如果坚持使用smart_open,只需在打开流时明确指定编码为utf-8即可解决非ASCII字符乱码:
代码示例
from smart_open import open import json def process_with_smart_open(bucket_name, s3_key): # 打开S3流时显式指定UTF-8编码 with open(f's3://{bucket_name}/{s3_key}', 'r', encoding='utf-8') as f: for line in f: try: json_obj = json.loads(line.strip()) # 替换为你的业务处理逻辑 # handle_json(json_obj) except json.JSONDecodeError as e: print(f"JSON解析失败: {e}") # 调用示例 process_with_smart_open('你的存储桶名称', '文件在S3中的路径/xxx.tar.gz')
乱码原因说明
smart_open默认可能使用系统默认编码(如ASCII)读取文本,遇到非ASCII字符时解码失败导致乱码。明确指定encoding='utf-8'后,会用UTF-8解码字节流,正确处理中文、emoji等非ASCII字符。
内容的提问来源于stack exchange,提问作者ciurlaro
相关产品推荐
相关产品推荐

