You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python在Cloud Function中实现GCS对象合成上传及报错排查

报错根因分析
  • 核心触发原因:调用list_blobs拉取桶内文件列表时,没有过滤待生成的目标文件feeds/file.csv,导致目标文件本身被加入合成源列表。你在合成前删除了旧的目标文件后,再调用bucket.get_blob()获取这个已被删除的源对象时会返回None,后续合成逻辑访问None.name属性时就抛出了AttributeError: 'NoneType' object has no attribute 'name'。
  • 代码中存在的其他会导致运行失败的问题:
    • 权限逻辑错误:原代码把合成生成的Blob对象当成Bucket实例传入IAM设置函数,权限配置完全不生效;且给整个存储桶配置公共读粒度过大,存在不必要的安全风险。
    • 缺失依赖导入:代码使用了List[str]类型注解但未导入typing.List,函数冷启动时会直接抛出导入错误。
    • 资源浪费:多个函数内重复初始化storage.Client()实例,会额外消耗Cloud Function的运行内存与执行时间。
    • 接口限制未适配:GCS的compose接口单次请求最多支持32个源文件,源文件数量超过32时原代码会直接报错。
    • 变量拼写错误:merge_files函数内先定义了拼写错误的desination变量,后续调用权限函数时使用正确拼写的destination会触发变量未定义错误。
    • 入口不符合规范:Cloud Function HTTP触发的入口函数必须接收request入参,原函数无该参数会导致触发失败。
可直接部署的修正版代码
import time
import datetime
from typing import List
from google.cloud import storage
# 如无自定义variables模块可删除下一行导入
# from variables import *
# import csv
# import json

MERCHANT_FILE_NAME = "/tmp/file.csv"
BUCKET_FILE_NAME = "feeds/file.csv"
BUCKET_NAME = "xxx-file-partial-feeds-bucket"
# GCS单次compose请求最大支持32个源文件
MAX_COMPOSE_PER_REQUEST = 32

def list_blobs(storage_client: storage.Client, bucket_name: str, exclude_name: str) -> List[str]:
    """拉取桶内所有blob名称,自动排除需要生成的目标文件"""
    blobs = storage_client.list_blobs(bucket_name)
    return [blob.name for blob in blobs if blob.name != exclude_name]

def compose_file(
    storage_client: storage.Client, 
    bucket_name: str, 
    blob_name_list: List[str], 
    destination_blob_name: str
) -> storage.Blob:
    """
    拼接多个源文件为目标文件
    合成前自动删除已存在的旧目标文件,自动适配32个源文件的接口限制
    """
    if not blob_name_list:
        raise ValueError("未找到任何可用于合成的源文件")
    
    bucket = storage_client.bucket(bucket_name)
    # 检查并删除已存在的旧目标文件
    destination = bucket.blob(destination_blob_name)
    if destination.exists():
        print(f"检测到已存在的旧目标文件{destination_blob_name},先执行删除")
        destination.delete()
        destination = bucket.blob(destination_blob_name)
    
    destination.content_type = "text/csv"

    # 源文件数量未超过接口限制时直接合成
    if len(blob_name_list) <= MAX_COMPOSE_PER_REQUEST:
        sources = [bucket.get_blob(blob_name) for blob_name in blob_name_list]
        # 过滤掉意外不存在的源文件
        sources = [blob for blob in sources if blob is not None]
        if not sources:
            raise RuntimeError("所有源文件均不存在,无法执行合成")
        destination.compose(sources)
    else:
        # 源文件超过32个时先分批合成临时对象
        temp_blobs = []
        batch_size = MAX_COMPOSE_PER_REQUEST - 1
        for idx, i in enumerate(range(0, len(blob_name_list), batch_size)):
            batch = blob_name_list[i:i+batch_size]
            temp_blob_name = f"temp/compose_part_{idx}_{int(time.time())}.csv"
            temp_blob = bucket.blob(temp_blob_name)
            temp_blob.content_type = "text/csv"
            batch_sources = [bucket.get_blob(name) for name in batch]
            batch_sources = [blob for blob in batch_sources if blob is not None]
            temp_blob.compose(batch_sources)
            temp_blobs.append(temp_blob)
        # 所有临时对象合成最终目标文件
        destination.compose(temp_blobs)
        # 清理临时分片文件
        for temp_blob in temp_blobs:
            temp_blob.delete()
    
    return destination

def set_blob_public(blob: storage.Blob):
    """仅为生成的目标文件设置公共读权限,不影响桶内其他文件"""
    blob.make_public()
    print(f"文件{blob.name}已设置公共读权限,公开访问地址:{blob.public_url}")

def merge_files(request):
    """Cloud Function HTTP触发入口"""
    try:
        storage_client = storage.Client()
        # 拉取源文件列表时排除目标文件本身
        blob_names = list_blobs(storage_client, BUCKET_NAME, BUCKET_FILE_NAME)
        print(f"共拉取到{len(blob_names)}个待合成源文件")
        destination_blob = compose_file(storage_client, BUCKET_NAME, blob_names, BUCKET_FILE_NAME)
        set_blob_public(destination_blob)
        return {"result": "success", "public_url": destination_blob.public_url}, 200
    
    except Exception as e:
        error_msg = f"合成任务执行失败:{str(e)}"
        print(error_msg)
        return {"error": error_msg}, 500
关键修正说明
  • 拉取文件列表时自动过滤目标合成文件,从根源避免把已删除的目标对象当成源文件读取的问题
  • 废弃整桶配置IAM的逻辑,直接调用Blob自带的make_public()方法为单个合成文件设置公共读,权限粒度最小,符合安全最佳实践
  • 补全缺失的类型导入,修复变量拼写错误,移除重复的Storage客户端初始化逻辑,减少不必要的资源消耗
  • 新增超过32个源文件时的分批合成逻辑,自动清理临时分片文件,完全适配GCS接口限制
  • 增加源文件为空、源文件不存在的异常判断,提前拦截无意义的API调用
  • 调整入口函数符合Cloud Function HTTP触发规范,执行成功后直接返回文件的公开访问地址
  • 保留了合成前删除旧目标文件的逻辑,避免旧版本文件残留影响合成结果

部署注意:需要给Cloud Function绑定的服务账号授予存储桶的storage.objects.list、storage.objects.delete、storage.objects.create、storage.objects.get权限,否则会出现权限不足的报错。

内容的提问来源于stack exchange,提问作者blob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 01:45:46