You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取S3桶多GZIP文件并合并转换为JSON格式

解决方案:S3 GZIP文件下载、合并与JSON格式转换

以下是基于Python实现的完整流程,假设你已通过boto3完成S3客户端的初始化:

步骤说明

  • 遍历指定的S3 GZIP文件列表
  • 下载并解压每个文件,将字节格式内容转换为字符串
  • 逐行解析JSON数据,合并所有内容并保存到本地

代码实现

import boto3
import gzip
import json
import os

# 初始化S3客户端(假设你已完成身份验证)
s3 = boto3.client('s3')
bucket_name = "你的S3桶名称"
file_list = ["file1.gz", "file2.gz", "file3.gz"]
local_save_dir = "./s3_merged_data"
merged_json_path = os.path.join(local_save_dir, "merged_data.json")

# 创建本地保存目录
os.makedirs(local_save_dir, exist_ok=True)

# 初始化合并后的JSON列表
merged_data = []

for file_name in file_list:
    # 临时本地文件路径
    local_gz_path = os.path.join(local_save_dir, file_name)
    
    # 从S3下载GZIP文件到本地
    s3.download_file(bucket_name, file_name, local_gz_path)
    
    # 解压并处理文件内容
    with gzip.open(local_gz_path, 'rb') as gz_file:
        # 读取所有字节并转换为字符串
        content_str = gz_file.read().decode('utf-8')
        # 按行分割内容(跳过空行)
        lines = [line.strip() for line in content_str.split('\n') if line.strip()]
        
        for line in lines:
            try:
                # 将单行字符串解析为JSON对象
                json_obj = json.loads(line)
                merged_data.append(json_obj)
            except json.JSONDecodeError as e:
                print(f"解析行失败: {line}, 错误信息: {e}")
                continue

# 将合并后的JSON数据保存到本地文件
with open(merged_json_path, 'w', encoding='utf-8') as f:
    json.dump(merged_data, f, ensure_ascii=False, indent=2)

print(f"合并完成,文件已保存至: {merged_json_path}")

注意事项

  • 替换代码中的bucket_name为你的实际S3桶名称
  • 如果文件体积较大,建议采用逐行写入的方式避免内存占用过高(示例为小文件场景的全量收集写法)
  • 示例中加入了JSON解析异常捕获,可根据业务需求调整错误处理逻辑

内容的提问来源于stack exchange,提问作者Lilly s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 20:25:22