Cloud Function跨项目存储桶复制文件内存超限问题排查与解决
问题:Cloud Function内存超限错误排查与解决
我正在开发一个Cloud Function,用于将文件从源存储桶复制到其他项目的存储桶中,源存储桶每日新增约500,000个文件。多次遇到以下错误:
Memory limit of 2048 MiB exceeded with 2050 MiB used. Consider increasing the memory limit
相关配置信息:
Trigger Type: Cloud Storage Event: finalizing/creating Memory Allocate: 2 GB timeout: 540 Maximum number of instance: 500
代码实现如下:
from google.cloud import storage from google.cloud.storage import Blob from datetime import date, datetime,timedelta def handle(event, context): try: file = event file_name = file['name'] trigger_time = file['timeCreated'] file_extention = file_name.split('/')[-1] plat_form = file_name.split('/')[1] is_jsondata = file_name.split('/')[3] device_id = file_name.split('/')[2] #active_date = datetime.now()+ timedelta(days=-1) active_date = trigger_time[0:10] print(f"trigger time is {active_date}") if plat_form == 'android' and is_jsondata == 'jsonData' and file_extention[0:10] == active_date: print(f"Processing Step1 {file_extention} and {file_extention[0:10]}") print(f"Processing File: {file['name']}.") storage_client = storage.Client(project='bucketA') #storage_client = storage.Client(project='bucketA') #storage_client_des = storage.Client(project='data-ingestion-production') source_bucket = storage_client.get_bucket('bucketa.appspot.com') #source_bucket = storage_client.get_bucket('bucketA.appspot.com') destination_bucket = storage_client.get_bucket('bucketb_tmp') #destination_object_name = storage_client.get_object('mobileDataSDK_dev/') prefix = 'movile/android/' blobs = list(source_bucket.list_blobs(prefix=prefix)) #target_name = file_name.split('/')[-1] #target_path = 'mobileDataSDK_dev/'+active_date.strftime("%Y%m%d")+'/'+device_id+'/' target_path = 'movile/android/'+active_date+ '/'+device_id+'/' for blob in blobs: if blob.name == file_name: source_blob = source_bucket.blob(blob.name) destination_path = target_path+file_extention #destination_path = target_path+blob.name new_blob = source_bucket.copy_blob( source_blob, destination_bucket, destination_path ) print(f"File move from {source_blob} to {new_blob}") else: print("Don't match criteria") except (RuntimeError, TypeError, NameError): print("Some errors") pass
错误原因分析
- 核心内存泄漏点:代码中执行
blobs = list(source_bucket.list_blobs(prefix=prefix)),会将movile/android/前缀下的所有Blob元数据一次性加载到内存。源存储桶每日新增50万文件,该前缀下的文件总量庞大,大量Blob对象直接耗尽2GB内存。 - 触发机制放大问题:Cloud Storage的
finalizing/creating事件会为每个新文件启动一个函数实例,短时间内大量文件创建会同时触发数百个实例,每个实例都加载全量Blob列表,进一步加剧内存压力。
解决方法
1. 移除冗余的全量Blob加载逻辑
当前代码加载所有Blob后再循环匹配当前触发文件的逻辑完全冗余,直接通过事件中的file_name获取目标Blob即可:
修改后的核心代码片段:
if plat_form == 'android' and is_jsondata == 'jsonData' and file_extention[0:10] == active_date: print(f"Processing Step1 {file_extention} and {file_extention[0:10]}") print(f"Processing File: {file['name']}.") storage_client = storage.Client(project='bucketA') source_bucket = storage_client.get_bucket('bucketa.appspot.com') destination_bucket = storage_client.get_bucket('bucketb_tmp') # 直接通过file_name获取源Blob,无需加载全量列表 source_blob = source_bucket.blob(file_name) target_path = 'movile/android/'+active_date+ '/'+device_id+'/' destination_path = target_path+file_extention new_blob = source_bucket.copy_blob( source_blob, destination_bucket, destination_path ) print(f"File copied from {source_blob} to {new_blob}")
2. 优化异常处理(辅助排查)
当前异常捕获范围过宽且日志模糊,建议针对性捕获并输出详细错误信息:
except Exception as e: print(f"Error processing file {file_name}: {str(e)}") raise # 重新抛出异常,便于Cloud Function记录完整错误日志
3. 内存配置调整(可选)
如果代码优化后仍存在内存波动,可适当提高内存配置,但优先代码优化(更高内存会增加运行成本)。
内容的提问来源于stack exchange,提问作者P.pp
相关产品推荐
相关产品推荐

