You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

S3不支持gzip.GzipFile上传?boto3上传gzip压缩文件报错如何解决

问题原因及解决方案

核心错误原因

你遇到的报错和S3、boto3是否支持gzip文件无关,完全是代码逻辑错误导致的,具体问题如下:

  • 缺失必要导入:既没有导入zipfile模块,也没有初始化s3_resource就直接调用,运行时会先触发这两类报错
  • 文件对象已关闭:gzip.open()的with代码块执行结束后,f_out会被自动关闭,此时将已关闭的GzipFile类型对象传给upload_fileobj,接口无法识别该类型,直接抛出对应错误
  • 无效代码冗余:gzip.compress(f_out)是完全错误的调用,该方法仅接收字节类型参数,无法直接处理文件对象,且生成的变量也没有被后续逻辑使用

修正方案

方案1:保留本地临时文件逻辑的修正版本

如果你需要在本地保留生成的gz文件,直接调用upload_file上传本地生成的文件即可,代码如下:

import boto3
import gzip
import shutil
import zipfile

# 补全s3 resource初始化
s3_resource = boto3.resource('s3')
bucket = s3_resource.Bucket('testunzipping')

with zipfile.ZipFile('/tmp/DataPump_10000838.zip', 'r') as zip_ref:
    # 先全量解压到临时目录
    zip_ref.extractall('/tmp/')
    testList = []
    for i in zip_ref.namelist():
        # 过滤MACOSX缓存文件和目录项
        if not i.startswith("__MACOSX/") and not i.endswith('/'):
            val = '/tmp/'+i
            testList.append(val)
            
    # 确认列表非空再删除首项,避免索引报错
    if testList:
        testList.pop(0)

    for i in testList:
        fileName = i.replace("/tmp/DataPump_10000838/", "") 
        fileName2 = i + '.gz'
        # 压缩生成本地gz文件
        with open(i, 'rb') as f_in:
            with gzip.open(fileName2, 'wb') as f_out:
                shutil.copyfileobj(f_in, f_out)
        # 直接上传本地gz文件
        bucket.upload_file(
            Filename=fileName2,
            Key=fileName + '.gz',
            ExtraArgs={'ContentType': "text/plain", 'ContentEncoding':'gzip'}
        )

方案2:内存压缩上传版本(推荐,适合Lambda等临时磁盘空间有限的环境)

不需要写入本地磁盘,直接在内存完成解压、压缩、上传全流程,代码如下:

import boto3
import gzip
import io
import zipfile

s3_resource = boto3.resource('s3')
bucket = s3_resource.Bucket('testunzipping')

with zipfile.ZipFile('/tmp/DataPump_10000838.zip', 'r') as zip_ref:
    testList = []
    for i in zip_ref.namelist():
        if not i.startswith("__MACOSX/") and not i.endswith('/'):
            testList.append(i)
            
    if testList:
        testList.pop(0)

    for file_key in testList:
        # 直接从原压缩包读取文件内容,无需写入本地
        with zip_ref.open(file_key, 'r') as f_in:
            file_content = f_in.read()
        
        # 内存中完成gzip压缩
        gz_buffer = io.BytesIO()
        with gzip.GzipFile(fileobj=gz_buffer, mode='wb') as f_out:
            f_out.write(file_content)
        # 重置缓冲区指针到开头,否则上传内容为空
        gz_buffer.seek(0)
        
        s3_key = file_key.replace("DataPump_10000838/", "")
        bucket.upload_fileobj(
            Fileobj=gz_buffer,
            Key=s3_key + '.gz',
            ExtraArgs={'ContentType': "text/plain", 'ContentEncoding':'gzip'}
        )

额外说明

你之前修改ContentType为application/gzip无效的原因是报错和HTTP头配置无关,仅和你传入的文件对象类型异常有关。S3完全支持任意格式的文件存储,gzip格式没有任何限制。

内容的提问来源于stack exchange,提问作者x89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 05:45:04