Python在Cloud Function中实现GCS对象合成上传及报错排查
报错根因分析
- 核心触发原因:调用
list_blobs拉取桶内文件列表时,没有过滤待生成的目标文件feeds/file.csv,导致目标文件本身被加入合成源列表。你在合成前删除了旧的目标文件后,再调用bucket.get_blob()获取这个已被删除的源对象时会返回None,后续合成逻辑访问None.name属性时就抛出了AttributeError: 'NoneType' object has no attribute 'name'。 - 代码中存在的其他会导致运行失败的问题:
- 权限逻辑错误:原代码把合成生成的Blob对象当成Bucket实例传入IAM设置函数,权限配置完全不生效;且给整个存储桶配置公共读粒度过大,存在不必要的安全风险。
- 缺失依赖导入:代码使用了
List[str]类型注解但未导入typing.List,函数冷启动时会直接抛出导入错误。 - 资源浪费:多个函数内重复初始化
storage.Client()实例,会额外消耗Cloud Function的运行内存与执行时间。 - 接口限制未适配:GCS的compose接口单次请求最多支持32个源文件,源文件数量超过32时原代码会直接报错。
- 变量拼写错误:
merge_files函数内先定义了拼写错误的desination变量,后续调用权限函数时使用正确拼写的destination会触发变量未定义错误。 - 入口不符合规范:Cloud Function HTTP触发的入口函数必须接收
request入参,原函数无该参数会导致触发失败。
可直接部署的修正版代码
import time import datetime from typing import List from google.cloud import storage # 如无自定义variables模块可删除下一行导入 # from variables import * # import csv # import json MERCHANT_FILE_NAME = "/tmp/file.csv" BUCKET_FILE_NAME = "feeds/file.csv" BUCKET_NAME = "xxx-file-partial-feeds-bucket" # GCS单次compose请求最大支持32个源文件 MAX_COMPOSE_PER_REQUEST = 32 def list_blobs(storage_client: storage.Client, bucket_name: str, exclude_name: str) -> List[str]: """拉取桶内所有blob名称,自动排除需要生成的目标文件""" blobs = storage_client.list_blobs(bucket_name) return [blob.name for blob in blobs if blob.name != exclude_name] def compose_file( storage_client: storage.Client, bucket_name: str, blob_name_list: List[str], destination_blob_name: str ) -> storage.Blob: """ 拼接多个源文件为目标文件 合成前自动删除已存在的旧目标文件,自动适配32个源文件的接口限制 """ if not blob_name_list: raise ValueError("未找到任何可用于合成的源文件") bucket = storage_client.bucket(bucket_name) # 检查并删除已存在的旧目标文件 destination = bucket.blob(destination_blob_name) if destination.exists(): print(f"检测到已存在的旧目标文件{destination_blob_name},先执行删除") destination.delete() destination = bucket.blob(destination_blob_name) destination.content_type = "text/csv" # 源文件数量未超过接口限制时直接合成 if len(blob_name_list) <= MAX_COMPOSE_PER_REQUEST: sources = [bucket.get_blob(blob_name) for blob_name in blob_name_list] # 过滤掉意外不存在的源文件 sources = [blob for blob in sources if blob is not None] if not sources: raise RuntimeError("所有源文件均不存在,无法执行合成") destination.compose(sources) else: # 源文件超过32个时先分批合成临时对象 temp_blobs = [] batch_size = MAX_COMPOSE_PER_REQUEST - 1 for idx, i in enumerate(range(0, len(blob_name_list), batch_size)): batch = blob_name_list[i:i+batch_size] temp_blob_name = f"temp/compose_part_{idx}_{int(time.time())}.csv" temp_blob = bucket.blob(temp_blob_name) temp_blob.content_type = "text/csv" batch_sources = [bucket.get_blob(name) for name in batch] batch_sources = [blob for blob in batch_sources if blob is not None] temp_blob.compose(batch_sources) temp_blobs.append(temp_blob) # 所有临时对象合成最终目标文件 destination.compose(temp_blobs) # 清理临时分片文件 for temp_blob in temp_blobs: temp_blob.delete() return destination def set_blob_public(blob: storage.Blob): """仅为生成的目标文件设置公共读权限,不影响桶内其他文件""" blob.make_public() print(f"文件{blob.name}已设置公共读权限,公开访问地址:{blob.public_url}") def merge_files(request): """Cloud Function HTTP触发入口""" try: storage_client = storage.Client() # 拉取源文件列表时排除目标文件本身 blob_names = list_blobs(storage_client, BUCKET_NAME, BUCKET_FILE_NAME) print(f"共拉取到{len(blob_names)}个待合成源文件") destination_blob = compose_file(storage_client, BUCKET_NAME, blob_names, BUCKET_FILE_NAME) set_blob_public(destination_blob) return {"result": "success", "public_url": destination_blob.public_url}, 200 except Exception as e: error_msg = f"合成任务执行失败:{str(e)}" print(error_msg) return {"error": error_msg}, 500
关键修正说明
- 拉取文件列表时自动过滤目标合成文件,从根源避免把已删除的目标对象当成源文件读取的问题
- 废弃整桶配置IAM的逻辑,直接调用Blob自带的
make_public()方法为单个合成文件设置公共读,权限粒度最小,符合安全最佳实践 - 补全缺失的类型导入,修复变量拼写错误,移除重复的Storage客户端初始化逻辑,减少不必要的资源消耗
- 新增超过32个源文件时的分批合成逻辑,自动清理临时分片文件,完全适配GCS接口限制
- 增加源文件为空、源文件不存在的异常判断,提前拦截无意义的API调用
- 调整入口函数符合Cloud Function HTTP触发规范,执行成功后直接返回文件的公开访问地址
- 保留了合成前删除旧目标文件的逻辑,避免旧版本文件残留影响合成结果
部署注意:需要给Cloud Function绑定的服务账号授予存储桶的
storage.objects.list、storage.objects.delete、storage.objects.create、storage.objects.get权限,否则会出现权限不足的报错。
内容的提问来源于stack exchange,提问作者blob
相关产品推荐
相关产品推荐

