GCP Storage存储桶下载文件为空无法读取内容解决方案
GCS文件下载与调用修复方案
核心问题根因
你遇到的空文件、后续调用失败问题由3个错误导致:
- 下载方法参数传反:
blob.download_blob_to_file第一个参数要求传入可写入的文件对象,原代码把blob自身作为第一个参数传入,根本没有把存储桶里的文件内容写入本地文件,所以生成空文件。 - 文件名不匹配:下载时本地保存的文件名是
new_file.py,后续subprocess调用时找的是file1_1.py,就算下载成功也会报文件不存在。 - 调试代码无效:打印blob长度时直接引用了
blob.download_as_string方法对象,没有加括号调用,根本拿不到实际文件大小,无法辅助排查问题。
第一步:修复GCS文件下载逻辑
直接用官方封装的download_to_filename方法,不需要手动管理文件句柄,避免参数顺序写错,同时加文件存在性校验:
from google.cloud import storage import os jsonkey = 'googl-cloudstorage-key.json' storage_client = storage.Client.from_service_account_json(jsonkey) def download_file_from_bucket(blob_name, file_path, bucket_name): print(f"download task: blob={blob_name}, local save path={file_path}") bucket = storage_client.get_bucket(bucket_name) print(f"connected bucket: {bucket.name}") blob = bucket.blob(blob_name) # 提前校验云端文件存在,避免下载不存在的空对象 if not blob.exists(): raise FileNotFoundError(f"Cloud file {blob_name} not exist in bucket {bucket_name}") print(f"cloud file size: {blob.size} bytes") # 直接下载到指定路径,二进制写入避免编码/换行符损坏脚本 blob.download_to_filename(file_path) print(f"local file saved, size: {os.path.getsize(file_path)} bytes") # 注意本地保存文件名和后续调用的文件名保持一致 download_file_from_bucket('file1.py', os.path.join(os.getcwd(),'file1_1.py'),'kk_bucket_1')
第二步:修复subprocess调用逻辑
统一文件路径,加异常捕获,避免环境不一致导致调用失败:
import subprocess import os import sys cust = ['cust1', 'cust2'] # 复用当前脚本的Python解释器路径,避免Airflow环境多Python版本冲突 PYTHON_PATH = sys.executable for c in cust: print(f"processing customer: {c}") target_script = os.path.join(os.getcwd(), 'file1_1.py') # 调用前校验脚本存在 if not os.path.exists(target_script): raise FileNotFoundError(f"Script {target_script} missing, check download step") print(f"calling script: {target_script}") exec_result = subprocess.run( [PYTHON_PATH, target_script, c], capture_output=True, text=True, check=True # 脚本执行报错时直接抛出异常,避免静默失败 ) print(f"script output:\n{exec_result.stdout}") if exec_result.stderr: print(f"script warning/error:\n{exec_result.stderr}")
Airflow部署注意事项
后续迁移到GCP Airflow(Cloud Composer)时做3个调整即可:
- 去掉本地服务账号密钥配置:把
storage_client = storage.Client.from_service_account_json(jsonkey)改成storage_client = storage.Client(),直接用Composer绑定的工作负载身份鉴权,不需要上传密钥文件,更安全。 - 下载路径改到临时目录:不要把脚本存在当前工作目录,Composer工作目录动态变化,建议存到
/tmp目录,比如下载路径改成os.path.join('/tmp', 'file1_1.py'),调用路径同步修改,避免权限问题。 - 权限校验:保证Composer服务账号有目标存储桶的
storage.objects.get权限,否则会报403错误。
内容的提问来源于stack exchange,提问作者Karan Alang
相关产品推荐
相关产品推荐

