AWS Lambda中S3 Zip文件解压后,对tmp目录文件Gzip压缩并上传至目标S3桶的方法咨询
1. Compressing Files Extracted to /tmp
If you're already extracting files to Lambda's /tmp directory, here's a straightforward way to compress each file to {filename}.csv.gz and upload to S3:
First, iterate over all files in /tmp, skip the original zip file and any directories, then use gzip.open to create compressed versions. This method uses efficient chunked copying to avoid loading entire files into memory:
import os import gzip import shutil import boto3 s3 = boto3.client('s3') tmp_dir = '/tmp' original_zip = '10000838.zip' target_bucket = 'your-target-bucket-name' # Iterate over all items in /tmp for filename in os.listdir(tmp_dir): file_path = os.path.join(tmp_dir, filename) # Skip directories and the original zip file if not os.path.isfile(file_path) or filename == original_zip: continue # Define path for compressed file (add .gz extension) compressed_path = f"{file_path}.gz" # Compress the file with open(file_path, 'rb') as input_file: with gzip.open(compressed_path, 'wb') as compressed_file: # Copy data in chunks to save memory shutil.copyfileobj(input_file, compressed_file) # Upload compressed file to S3 s3.upload_file(compressed_path, target_bucket, f"{filename}.gz") # Optional: Clean up temporary files to free space os.remove(file_path) os.remove(compressed_path)
Key Notes:
gzip.open(compressed_path, 'wb')creates the compressed file at the target path you specify (here,/tmp/{filename}.gz). The first argument is where you want to save the compressed output.shutil.copyfileobjis preferred over reading the entire file into memory because it processes data in small chunks, which is critical for Lambda's memory limits.- Saving to
/tmpis ideal because it's Lambda's writable temporary storage, ands3.upload_filecan directly read from this path.
Regarding the compress function: It's designed for in-memory compression of bytes objects, which isn't necessary here. Using gzip.open is more efficient for file-based operations.
2. Streaming Approach (No Temporary Files)
For better efficiency (saving disk space and time), you can stream data directly from the zip file, compress it on the fly, and upload to S3 without ever writing files to /tmp. Here's how to implement this:
import zipfile import gzip import shutil import io import boto3 s3 = boto3.client('s3') target_bucket = 'your-target-bucket-name' with zipfile.ZipFile('/tmp/10000838.zip', 'r') as zip_ref: # Filter out __MACOSX files and directories valid_items = [ item for item in zip_ref.namelist() if not item.startswith("__MACOSX/") and not item.endswith('/') ] for item in valid_items: # Define S3 key for the compressed file (e.g., "data.csv.gz") s3_key = f"{item}.gz" # Open the zip member as a binary stream with zip_ref.open(item, 'rb') as zip_stream: # Use a BytesIO buffer to hold compressed data temporarily compressed_buffer = io.BytesIO() with gzip.GzipFile(fileobj=compressed_buffer, mode='wb') as gz_file: shutil.copyfileobj(zip_stream, gz_file) # Reset buffer to the start for uploading compressed_buffer.seek(0) # Upload directly to S3 s3.upload_fileobj(compressed_buffer, target_bucket, s3_key)
Advanced: True Streaming (No Buffer)
If you're dealing with extremely large files and want to avoid even the temporary buffer, use a custom stream wrapper that compresses chunks on the fly:
# Add this class before the upload loop class StreamingGzipCompressor: def __init__(self, source_stream): self.source = source_stream self.gz_buffer = io.BytesIO() self.gzip_file = gzip.GzipFile(fileobj=self.gz_buffer, mode='wb') self.pending_data = b'' def read(self, size=-1): while not self.pending_data and not self.source.closed: # Read chunk from zip stream chunk = self.source.read(16384) if not chunk: # Finalize compression self.gzip_file.close() self.pending_data = self.gz_buffer.getvalue() self.gz_buffer.close() break # Compress the chunk self.gzip_file.write(chunk) self.gz_buffer.seek(0) self.pending_data = self.gz_buffer.read() self.gz_buffer.seek(0) self.gz_buffer.truncate() # Return requested amount of data output, self.pending_data = self.pending_data[:size], self.pending_data[size:] return output # Use it in the loop instead of the BytesIO approach with zip_ref.open(item, 'rb') as zip_stream: s3.upload_fileobj(StreamingGzipCompressor(zip_stream), target_bucket, s3_key)
Key Notes:
- This approach skips extracting files to
/tmpentirely, reducing disk usage and latency. zip_ref.open(item, 'rb')gives direct access to the zip member's content as a stream.- The custom
StreamingGzipCompressorfeeds compressed chunks to S3 incrementally, making it ideal for large files that might exceed Lambda's memory limits.
内容的提问来源于stack exchange,提问作者x89

