You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GCP Storage存储桶下载文件为空无法读取内容解决方案

GCS文件下载与调用修复方案

核心问题根因

你遇到的空文件、后续调用失败问题由3个错误导致:

  • 下载方法参数传反:blob.download_blob_to_file 第一个参数要求传入可写入的文件对象,原代码把blob自身作为第一个参数传入,根本没有把存储桶里的文件内容写入本地文件,所以生成空文件。
  • 文件名不匹配:下载时本地保存的文件名是new_file.py,后续subprocess调用时找的是file1_1.py,就算下载成功也会报文件不存在。
  • 调试代码无效:打印blob长度时直接引用了blob.download_as_string方法对象,没有加括号调用,根本拿不到实际文件大小,无法辅助排查问题。

第一步:修复GCS文件下载逻辑

直接用官方封装的download_to_filename方法,不需要手动管理文件句柄,避免参数顺序写错,同时加文件存在性校验:

from google.cloud import storage
import os

jsonkey = 'googl-cloudstorage-key.json'
storage_client = storage.Client.from_service_account_json(jsonkey)

def download_file_from_bucket(blob_name, file_path, bucket_name):
    print(f"download task: blob={blob_name}, local save path={file_path}")
    bucket = storage_client.get_bucket(bucket_name)
    print(f"connected bucket: {bucket.name}")
    blob = bucket.blob(blob_name)
    
    # 提前校验云端文件存在,避免下载不存在的空对象
    if not blob.exists():
        raise FileNotFoundError(f"Cloud file {blob_name} not exist in bucket {bucket_name}")
    print(f"cloud file size: {blob.size} bytes")
    
    # 直接下载到指定路径,二进制写入避免编码/换行符损坏脚本
    blob.download_to_filename(file_path)
    print(f"local file saved, size: {os.path.getsize(file_path)} bytes")

# 注意本地保存文件名和后续调用的文件名保持一致
download_file_from_bucket('file1.py', os.path.join(os.getcwd(),'file1_1.py'),'kk_bucket_1')

第二步:修复subprocess调用逻辑

统一文件路径,加异常捕获,避免环境不一致导致调用失败:

import subprocess
import os
import sys

cust = ['cust1', 'cust2']
# 复用当前脚本的Python解释器路径,避免Airflow环境多Python版本冲突
PYTHON_PATH = sys.executable

for c in cust:
    print(f"processing customer: {c}")
    target_script = os.path.join(os.getcwd(), 'file1_1.py')
    # 调用前校验脚本存在
    if not os.path.exists(target_script):
        raise FileNotFoundError(f"Script {target_script} missing, check download step")
    print(f"calling script: {target_script}")
    
    exec_result = subprocess.run(
        [PYTHON_PATH, target_script, c],
        capture_output=True,
        text=True,
        check=True # 脚本执行报错时直接抛出异常,避免静默失败
    )
    print(f"script output:\n{exec_result.stdout}")
    if exec_result.stderr:
        print(f"script warning/error:\n{exec_result.stderr}")

Airflow部署注意事项

后续迁移到GCP Airflow(Cloud Composer)时做3个调整即可:

  • 去掉本地服务账号密钥配置:把storage_client = storage.Client.from_service_account_json(jsonkey)改成storage_client = storage.Client(),直接用Composer绑定的工作负载身份鉴权,不需要上传密钥文件,更安全。
  • 下载路径改到临时目录:不要把脚本存在当前工作目录,Composer工作目录动态变化,建议存到/tmp目录,比如下载路径改成os.path.join('/tmp', 'file1_1.py'),调用路径同步修改,避免权限问题。
  • 权限校验:保证Composer服务账号有目标存储桶的storage.objects.get权限,否则会报403错误。

内容的提问来源于stack exchange,提问作者Karan Alang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 14:15:33