You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Cloud Storage Python客户端上传至GCS Bucket时丢失记录

GCS upload_from_filename上传CSV出现随机记录丢失,但本地文件完整(Vertex AI流水线场景)

问题背景

运行Vertex AI流水线,组件逻辑如下:

  • 拥有两个GCS Bucket:归档桶、上传桶
  • 流水线生成两个预测CSV文件file_1和file_2
  • 代码先检查上传桶是否存在文件,若存在则归档至归档桶并删除,随后上传新生成的文件

归档删除步骤执行正常,但使用upload_from_filename方法上传后,CSV文件出现随机记录丢失;作为Vertex AI组件工件输出的本地文件记录完整。

组件代码

from google.cloud import storage
import time
from datetime import datetime

file_1_name = '<insert csv file 1 name>'
file_2_name = '<insert csv file 2 name>'

archive_bucket = storage_client.bucket(archive_bucket_name)
upload_bucket = storage_client.bucket(target_bucket_name)

# 生成上传桶中的blob列表
target_bucket_blobs = list(upload_bucket.list_blobs())

# 若上传桶无文件,直接跳过归档;若有文件则归档并删除
if target_bucket_blobs:
    for blob in target_bucket_blobs:
        upload_bucket.copy_blob(blob, archive_bucket, blob.name)
        blob.delete()

# 等待5分钟(原代码逻辑)
time.sleep(300)

# 上传新预测文件
current_time = datetime.today().strftime("%Y-%m-%d - %Hh-%Mm-%Ss")
print(f'Uploading predictions to {target_bucket_name}.')

file_1_blob = upload_bucket.blob(file_1_name + '_' + current_time)
file_2_blob = upload_bucket.blob(file_2_name + '_' + current_time)

file_1_blob.upload_from_filename(file_1_path, content_type='text/csv')
file_2_blob.upload_from_filename(file_2_path, content_type='text/csv')

排查与解决方案

1. 确保本地文件写入完成

随机丢记录最常见的原因是流水线生成文件后,文件尚未完全写入磁盘就触发上传。upload_from_filename直接读取文件路径,若文件仍在被写入(比如流水线的前序步骤未完全关闭文件句柄),会导致读取不完整。

修复方式:改用upload_from_file,通过文件对象读取,确保文件已完全写入并关闭:

# 替换原上传代码块
print(f'Uploading predictions to {target_bucket_name}.')
current_time = datetime.today().strftime("%Y-%m-%d - %Hh-%Mm-%Ss")

# 以二进制只读模式打开文件,确保文件已写入完成
with open(file_1_path, 'rb') as f1, open(file_2_path, 'rb') as f2:
    file_1_blob = upload_bucket.blob(f"{file_1_name}_{current_time}")
    file_1_blob.upload_from_file(f1, content_type='text/csv')
    
    file_2_blob = upload_bucket.blob(f"{file_2_name}_{current_time}")
    file_2_blob.upload_from_file(f2, content_type='text/csv')

2. 移除不必要的time.sleep(300)

copy_blob和blob.delete()都是同步操作,执行完成后归档和删除已彻底完成,无需等待。这个sleep不仅解决不了问题,还会浪费流水线运行时间,直接移除即可。

3. 校验文件上传完整性

通过对比本地文件和GCS文件的哈希值,确认是否是上传过程导致的数据丢失:

import hashlib

def calculate_md5(file_path):
    hash_md5 = hashlib.md5()
    with open(file_path, 'rb') as f:
        for chunk in iter(lambda: f.read(4096), b''):
            hash_md5.update(chunk)
    return hash_md5.hexdigest()

# 上传后校验
local_md5_1 = calculate_md5(file_1_path)
gcs_md5_1 = file_1_blob.md5_hash
if local_md5_1 != gcs_md5_1:
    raise Exception(f"File 1上传失败:本地MD5 {local_md5_1} != GCS MD5 {gcs_md5_1}")

local_md5_2 = calculate_md5(file_2_path)
gcs_md5_2 = file_2_blob.md5_hash
if local_md5_2 != gcs_md5_2:
    raise Exception(f"File 2上传失败:本地MD5 {local_md5_2} != GCS MD5 {gcs_md5_2}")

4. 更新google-cloud-storage库

旧版本的客户端库可能存在文件上传的bug,更新到最新稳定版:

pip install --upgrade google-cloud-storage

5. 检查Vertex AI组件的文件权限

确保组件运行的服务账号对本地文件有完整的读取权限,避免因权限问题导致读取文件时截断内容。


内容的提问来源于stack exchange,提问作者AndrewJaeyoung

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 21:18:23