You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Lambda中S3 Zip文件解压后,对tmp目录文件Gzip压缩并上传至目标S3桶的方法咨询

Solution for Compressing Files from S3 Zip and Uploading to Another Bucket

1. Compressing Files Extracted to /tmp

If you're already extracting files to Lambda's /tmp directory, here's a straightforward way to compress each file to {filename}.csv.gz and upload to S3:

First, iterate over all files in /tmp, skip the original zip file and any directories, then use gzip.open to create compressed versions. This method uses efficient chunked copying to avoid loading entire files into memory:

import os
import gzip
import shutil
import boto3

s3 = boto3.client('s3')
tmp_dir = '/tmp'
original_zip = '10000838.zip'
target_bucket = 'your-target-bucket-name'

# Iterate over all items in /tmp
for filename in os.listdir(tmp_dir):
    file_path = os.path.join(tmp_dir, filename)
    
    # Skip directories and the original zip file
    if not os.path.isfile(file_path) or filename == original_zip:
        continue
    
    # Define path for compressed file (add .gz extension)
    compressed_path = f"{file_path}.gz"
    
    # Compress the file
    with open(file_path, 'rb') as input_file:
        with gzip.open(compressed_path, 'wb') as compressed_file:
            # Copy data in chunks to save memory
            shutil.copyfileobj(input_file, compressed_file)
    
    # Upload compressed file to S3
    s3.upload_file(compressed_path, target_bucket, f"{filename}.gz")
    
    # Optional: Clean up temporary files to free space
    os.remove(file_path)
    os.remove(compressed_path)

Key Notes:

  • gzip.open(compressed_path, 'wb') creates the compressed file at the target path you specify (here, /tmp/{filename}.gz). The first argument is where you want to save the compressed output.
  • shutil.copyfileobj is preferred over reading the entire file into memory because it processes data in small chunks, which is critical for Lambda's memory limits.
  • Saving to /tmp is ideal because it's Lambda's writable temporary storage, and s3.upload_file can directly read from this path.

Regarding the compress function: It's designed for in-memory compression of bytes objects, which isn't necessary here. Using gzip.open is more efficient for file-based operations.


2. Streaming Approach (No Temporary Files)

For better efficiency (saving disk space and time), you can stream data directly from the zip file, compress it on the fly, and upload to S3 without ever writing files to /tmp. Here's how to implement this:

import zipfile
import gzip
import shutil
import io
import boto3

s3 = boto3.client('s3')
target_bucket = 'your-target-bucket-name'

with zipfile.ZipFile('/tmp/10000838.zip', 'r') as zip_ref:
    # Filter out __MACOSX files and directories
    valid_items = [
        item for item in zip_ref.namelist() 
        if not item.startswith("__MACOSX/") and not item.endswith('/')
    ]
    
    for item in valid_items:
        # Define S3 key for the compressed file (e.g., "data.csv.gz")
        s3_key = f"{item}.gz"
        
        # Open the zip member as a binary stream
        with zip_ref.open(item, 'rb') as zip_stream:
            # Use a BytesIO buffer to hold compressed data temporarily
            compressed_buffer = io.BytesIO()
            with gzip.GzipFile(fileobj=compressed_buffer, mode='wb') as gz_file:
                shutil.copyfileobj(zip_stream, gz_file)
            
            # Reset buffer to the start for uploading
            compressed_buffer.seek(0)
            
            # Upload directly to S3
            s3.upload_fileobj(compressed_buffer, target_bucket, s3_key)

Advanced: True Streaming (No Buffer)

If you're dealing with extremely large files and want to avoid even the temporary buffer, use a custom stream wrapper that compresses chunks on the fly:

# Add this class before the upload loop
class StreamingGzipCompressor:
    def __init__(self, source_stream):
        self.source = source_stream
        self.gz_buffer = io.BytesIO()
        self.gzip_file = gzip.GzipFile(fileobj=self.gz_buffer, mode='wb')
        self.pending_data = b''
    
    def read(self, size=-1):
        while not self.pending_data and not self.source.closed:
            # Read chunk from zip stream
            chunk = self.source.read(16384)
            if not chunk:
                # Finalize compression
                self.gzip_file.close()
                self.pending_data = self.gz_buffer.getvalue()
                self.gz_buffer.close()
                break
            
            # Compress the chunk
            self.gzip_file.write(chunk)
            self.gz_buffer.seek(0)
            self.pending_data = self.gz_buffer.read()
            self.gz_buffer.seek(0)
            self.gz_buffer.truncate()
        
        # Return requested amount of data
        output, self.pending_data = self.pending_data[:size], self.pending_data[size:]
        return output

# Use it in the loop instead of the BytesIO approach
with zip_ref.open(item, 'rb') as zip_stream:
    s3.upload_fileobj(StreamingGzipCompressor(zip_stream), target_bucket, s3_key)

Key Notes:

  • This approach skips extracting files to /tmp entirely, reducing disk usage and latency.
  • zip_ref.open(item, 'rb') gives direct access to the zip member's content as a stream.
  • The custom StreamingGzipCompressor feeds compressed chunks to S3 incrementally, making it ideal for large files that might exceed Lambda's memory limits.

内容的提问来源于stack exchange,提问作者x89

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 19:59:08