You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需下载,在Google Cloud Storage中压缩多文件的GSUtil方法咨询

Can gsutil directly zip multiple GCS files into a single ZIP without downloading them?

Great question! Unfortunately, gsutil doesn't have a native command that lets you directly compress multiple files in a Google Cloud Storage (GCS) bucket into a ZIP archive without first downloading them to a local or cloud-based filesystem. But don't worry—there are reliable workarounds that let you achieve this entirely in the cloud, avoiding local downloads.

The most efficient way to handle this is to deploy a small cloud function that interacts directly with GCS. Here's a simplified Python example using the google-cloud-storage library:

import zipfile
from io import BytesIO
from google.cloud import storage

def zip_gcs_files(event, context):
    # Configure your bucket and target files
    source_bucket_name = "your-source-bucket"
    target_bucket_name = "your-target-bucket"
    files_to_zip = ["path/to/file1.txt", "path/to/image.jpg", "docs/report.pdf"]
    zip_filename = "compressed_files.zip"

    storage_client = storage.Client()
    source_bucket = storage_client.bucket(source_bucket_name)
    target_bucket = storage_client.bucket(target_bucket_name)

    # Create an in-memory ZIP file
    zip_buffer = BytesIO()
    with zipfile.ZipFile(zip_buffer, "w", zipfile.ZIP_DEFLATED) as zip_file:
        for file_path in files_to_zip:
            blob = source_bucket.blob(file_path)
            # Read file content directly from GCS without local download
            file_content = blob.download_as_bytes()
            # Add the file to the ZIP, preserving original path/filename
            zip_file.writestr(file_path, file_content)
    
    # Reset buffer position to start for upload
    zip_buffer.seek(0)
    # Upload the finished ZIP to your target bucket
    target_blob = target_bucket.blob(zip_filename)
    target_blob.upload_from_file(zip_buffer, content_type="application/zip")

You can trigger this function manually via HTTP, set it to run on a schedule, or tie it to bucket events (like new file uploads). Just ensure the service account linked to the function has storage.objects.get permissions on the source bucket and storage.objects.create permissions on the target bucket.

Alternative: Use a Temporary Cloud VM (for large files)

If you're working with very large files that might exceed Cloud Functions' memory limits, spin up a temporary Compute Engine VM:

  • SSH into the VM using gcloud compute ssh [vm-name]
  • Confirm gsutil and zip are installed (they’re preloaded on Google’s base OS images)
  • Use gsutil cp to copy target files to the VM’s temporary filesystem (or /tmp for in-memory handling)
  • Compress the files with zip compressed_files.zip file1.txt file2.jpg ...
  • Upload the ZIP back to GCS with gsutil cp compressed_files.zip gs://your-bucket/
  • Delete the VM immediately after to avoid unnecessary costs

Important Caveat About Pipe Methods

You might come across suggestions for pipe-based commands like gsutil cat gs://bucket/file* | zip -@ gs://bucket/output.zip, but this won’t work as intended. The zip -@ flag reads filenames (not file content) from stdin, so this approach will fail to structure the ZIP archive with individual, accessible files. Stick to the cloud function or VM methods for reliable results.

内容的提问来源于stack exchange,提问作者Kathan Shah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:02:54