Python处理S3解压文件gzip上传时报Negative seek in write mode错误如何修复
错误原因
报错的核心原因是处于wb写入模式的gzip.GzipFile对象不支持随机寻址(seek)操作,你错误地对gzip写入对象调用了seek(0),实际应该操作的是存储压缩后数据的底层BytesIO对象。
修复步骤
- 写完压缩数据后,先关闭gzip对象,确保所有压缩内容都被完整写入底层的
BytesIO,避免数据丢失 - 改为对
destination_file(存储压缩数据的BytesIO实例)调用seek(0),将文件指针移到起始位置 - 同步修复代码中两个额外问题:1. 未定义的变量
fileName,改为你预先定义的s3_filename(如果需要按原文件名保存,也可以取路径的最后一段生成);2. for循环缩进错误,需要和上层变量定义对齐
修复后的完整代码
import boto3 import contextlib import gzip import io s3_filename = 'sample' s3_resource = boto3.resource('s3') bucket = s3_resource.Bucket('testunzipping') # 修复缩进问题 for i in testList: with contextlib.ExitStack() as stack: source_file = stack.enter_context(open(i , mode="rb")) destination_file = io.BytesIO() destination_file_gz = stack.enter_context(gzip.GzipFile(fileobj=destination_file, mode='wb')) while True: chunk = source_file.read(1024) if not chunk: break destination_file_gz.write(chunk) # 退出ExitStack后gzip对象会自动关闭,所有数据已写入destination_file destination_file.seek(0) # 修复未定义的fileName变量,按需调整命名规则 # 如果需要保留原文件名加.gz后缀,可以改为:f"{i.split('/')[-1]}.gz" bucket.upload_fileobj(destination_file, f"{s3_filename}_{i.split('/')[-1]}.gz")
内容的提问来源于stack exchange,提问作者x89
相关产品推荐
相关产品推荐

