无需下载,在Google Cloud Storage中压缩多文件的GSUtil方法咨询
Great question! Unfortunately, gsutil doesn't have a native command that lets you directly compress multiple files in a Google Cloud Storage (GCS) bucket into a ZIP archive without first downloading them to a local or cloud-based filesystem. But don't worry—there are reliable workarounds that let you achieve this entirely in the cloud, avoiding local downloads.
Recommended Workaround: Use Cloud Functions or Cloud Run
The most efficient way to handle this is to deploy a small cloud function that interacts directly with GCS. Here's a simplified Python example using the google-cloud-storage library:
import zipfile from io import BytesIO from google.cloud import storage def zip_gcs_files(event, context): # Configure your bucket and target files source_bucket_name = "your-source-bucket" target_bucket_name = "your-target-bucket" files_to_zip = ["path/to/file1.txt", "path/to/image.jpg", "docs/report.pdf"] zip_filename = "compressed_files.zip" storage_client = storage.Client() source_bucket = storage_client.bucket(source_bucket_name) target_bucket = storage_client.bucket(target_bucket_name) # Create an in-memory ZIP file zip_buffer = BytesIO() with zipfile.ZipFile(zip_buffer, "w", zipfile.ZIP_DEFLATED) as zip_file: for file_path in files_to_zip: blob = source_bucket.blob(file_path) # Read file content directly from GCS without local download file_content = blob.download_as_bytes() # Add the file to the ZIP, preserving original path/filename zip_file.writestr(file_path, file_content) # Reset buffer position to start for upload zip_buffer.seek(0) # Upload the finished ZIP to your target bucket target_blob = target_bucket.blob(zip_filename) target_blob.upload_from_file(zip_buffer, content_type="application/zip")
You can trigger this function manually via HTTP, set it to run on a schedule, or tie it to bucket events (like new file uploads). Just ensure the service account linked to the function has storage.objects.get permissions on the source bucket and storage.objects.create permissions on the target bucket.
Alternative: Use a Temporary Cloud VM (for large files)
If you're working with very large files that might exceed Cloud Functions' memory limits, spin up a temporary Compute Engine VM:
- SSH into the VM using
gcloud compute ssh [vm-name] - Confirm
gsutilandzipare installed (they’re preloaded on Google’s base OS images) - Use
gsutil cpto copy target files to the VM’s temporary filesystem (or/tmpfor in-memory handling) - Compress the files with
zip compressed_files.zip file1.txt file2.jpg ... - Upload the ZIP back to GCS with
gsutil cp compressed_files.zip gs://your-bucket/ - Delete the VM immediately after to avoid unnecessary costs
Important Caveat About Pipe Methods
You might come across suggestions for pipe-based commands like gsutil cat gs://bucket/file* | zip -@ gs://bucket/output.zip, but this won’t work as intended. The zip -@ flag reads filenames (not file content) from stdin, so this approach will fail to structure the ZIP archive with individual, accessible files. Stick to the cloud function or VM methods for reliable results.
内容的提问来源于stack exchange,提问作者Kathan Shah

