You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python SDK批量删除Google Cloud Storage指定子目录

如何通过Python SDK批量删除GCS中匹配通配符路径的文件?

问题描述

现有Python代码:

storage_client = storage.Client()
bucket = storage.Bucket(storage_client, name="mybucket")
blobs = storage_client.list_blobs(bucket_or_name=bucket, prefix="content/")
print('Blobs:')
for blob in blobs:
   print(blob.name)

存储层级结构:

  • content/{uid}/subContentDirectory1/anyFiles*
  • content/{uid}/subContentDirectory2/anyFiles*

目标:删除所有content/{uid}/subContentDirectory1目录下的所有文件({uid}为通配符,可能匹配数百万条结果)。当前代码可获取content下所有子目录的文件,但不知道如何筛选并仅删除指定子目录的内容。由于代码将运行在AWS Lambda中,更倾向于使用Python SDK而非gsutil,求解决方案。

解决方案

1. 筛选并删除目标Blob

GCS没有真正的“目录”概念,所有对象都是以路径形式存储的Blob。以下两种方法可精准定位并删除目标路径下的文件:

方法一:前缀遍历+路径匹配

直接遍历content/前缀下的所有Blob,通过路径字符串判断筛选出subContentDirectory1下的文件:

from google.cloud import storage

def delete_target_blobs():
    storage_client = storage.Client()
    bucket_name = "mybucket"
    bucket = storage_client.bucket(bucket_name)
    
    # 遍历content前缀下的所有Blob
    blobs = storage_client.list_blobs(bucket, prefix="content/")
    
    # 筛选包含subContentDirectory1路径的Blob,分批删除
    target_blobs = []
    for blob in blobs:
        if "/subContentDirectory1/" in blob.name:
            target_blobs.append(blob)
            # 每积累1000个就批量删除,避免内存占用过高
            if len(target_blobs) == 1000:
                bucket.delete_blobs(target_blobs)
                print(f"已删除1000个文件")
                target_blobs = []
    
    # 删除剩余的Blob
    if target_blobs:
        bucket.delete_blobs(target_blobs)
        print(f"已删除剩余{len(target_blobs)}个文件")

方法二:层级递归匹配

先获取所有content/{uid}/层级的路径,再拼接subContentDirectory1前缀批量删除,这种方式更高效,避免遍历无关Blob:

from google.cloud import storage

def delete_target_blobs():
    storage_client = storage.Client()
    bucket_name = "mybucket"
    bucket = storage_client.bucket(bucket_name)
    
    # 获取content/下的所有一级子目录(即content/{uid}/)
    prefix = "content/"
    delimiter = "/"
    results = storage_client.list_blobs(bucket, prefix=prefix, delimiter=delimiter)
    
    # 遍历每个uid目录,处理对应的subContentDirectory1路径
    for uid_prefix in results.prefixes:
        target_prefix = f"{uid_prefix}subContentDirectory1/"
        # 批量获取该前缀下的所有Blob并分批删除
        target_blobs = []
        for blob in storage_client.list_blobs(bucket, prefix=target_prefix):
            target_blobs.append(blob)
            if len(target_blobs) == 1000:
                bucket.delete_blobs(target_blobs)
                print(f"已删除路径{target_prefix}下的1000个文件")
                target_blobs = []
        if target_blobs:
            bucket.delete_blobs(target_blobs)
            print(f"已删除路径{target_prefix}下剩余{len(target_blobs)}个文件")

2. 百万级Blob的优化建议

  • 分批删除:GCS的delete_blobs接口单次最多支持1000个Blob,代码中加入批量积累逻辑,避免单次请求过大。
  • Lambda适配:AWS Lambda默认执行时间上限15分钟,若Blob数量极大:
    • 提升Lambda内存配置(更高内存对应更高CPU和网络带宽,加快遍历和删除速度)
    • 采用异步拆分任务,比如通过SQS将不同uid的删除任务拆分,分多次执行
  • 权限配置:确保Lambda的执行角色拥有GCS的storage.objects.delete权限,避免删除失败。

内容的提问来源于stack exchange,提问作者Tom3652

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 08:59:16