You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python boto3流式读取S3中的.tar.gz压缩JSON文件?

解决方案:流式处理S3上的大体积.tar.gz JSON文件

一、原生Python实现(无第三方依赖)

直接用boto3+tarfile+json实现流式处理,彻底规避内存溢出和EOF错误:

实现逻辑

  1. 通过boto3获取S3对象的流式响应体,不下载整个文件到本地
  2. 用tarfile的流式模式(r|gz)解压,分块处理压缩数据,不占用大量内存
  3. 遍历压缩包内的每个文件,逐行读取JSON对象并处理

代码示例

import boto3
import tarfile
import json

def process_s3_targz_json(bucket_name, s3_key):
    s3 = boto3.client('s3')
    # 获取流式响应,避免一次性下载大文件
    response = s3.get_object(Bucket=bucket_name, Key=s3_key)
    stream = response['Body']

    # 流式解压tar.gz,r|gz模式支持边读边解
    with tarfile.open(fileobj=stream, mode='r|gz') as tar:
        for member in tar:
            if member.isfile():
                file_obj = tar.extractfile(member)
                if not file_obj:
                    continue
                # 逐行读取并解析JSON
                for line in file_obj:
                    try:
                        # 显式用UTF-8解码,处理非ASCII字符
                        json_obj = json.loads(line.decode('utf-8').strip())
                        # 替换为你的业务处理逻辑
                        # handle_json(json_obj)
                    except json.JSONDecodeError as e:
                        print(f"JSON解析失败: {e}")
                    except UnicodeDecodeError as e:
                        print(f"字符解码失败: {e}")
                file_obj.close()

# 调用示例
process_s3_targz_json('你的存储桶名称', '文件在S3中的路径/xxx.tar.gz')

EOFError解决说明

之前的内存解压方式是一次性把整个压缩文件加载到内存再解压,30GB的解压后文件远超内存容量,导致解压过程中数据截断或内存耗尽,触发EOFError。流式处理是分块读取、分块解压,每块仅占用少量内存,完全避免这个问题。


二、修复smart_open的乱码问题

如果坚持使用smart_open,只需在打开流时明确指定编码为utf-8即可解决非ASCII字符乱码:

代码示例

from smart_open import open
import json

def process_with_smart_open(bucket_name, s3_key):
    # 打开S3流时显式指定UTF-8编码
    with open(f's3://{bucket_name}/{s3_key}', 'r', encoding='utf-8') as f:
        for line in f:
            try:
                json_obj = json.loads(line.strip())
                # 替换为你的业务处理逻辑
                # handle_json(json_obj)
            except json.JSONDecodeError as e:
                print(f"JSON解析失败: {e}")

# 调用示例
process_with_smart_open('你的存储桶名称', '文件在S3中的路径/xxx.tar.gz')

乱码原因说明

smart_open默认可能使用系统默认编码(如ASCII)读取文本,遇到非ASCII字符时解码失败导致乱码。明确指定encoding='utf-8'后,会用UTF-8解码字节流,正确处理中文、emoji等非ASCII字符。


内容的提问来源于stack exchange,提问作者ciurlaro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 02:30:53