使用Cloud Storage Python客户端上传至GCS Bucket时丢失记录
GCS upload_from_filename上传CSV出现随机记录丢失,但本地文件完整(Vertex AI流水线场景)
问题背景
运行Vertex AI流水线,组件逻辑如下:
- 拥有两个GCS Bucket:归档桶、上传桶
- 流水线生成两个预测CSV文件
file_1和file_2 - 代码先检查上传桶是否存在文件,若存在则归档至归档桶并删除,随后上传新生成的文件
归档删除步骤执行正常,但使用upload_from_filename方法上传后,CSV文件出现随机记录丢失;作为Vertex AI组件工件输出的本地文件记录完整。
组件代码
from google.cloud import storage import time from datetime import datetime file_1_name = '<insert csv file 1 name>' file_2_name = '<insert csv file 2 name>' archive_bucket = storage_client.bucket(archive_bucket_name) upload_bucket = storage_client.bucket(target_bucket_name) # 生成上传桶中的blob列表 target_bucket_blobs = list(upload_bucket.list_blobs()) # 若上传桶无文件,直接跳过归档;若有文件则归档并删除 if target_bucket_blobs: for blob in target_bucket_blobs: upload_bucket.copy_blob(blob, archive_bucket, blob.name) blob.delete() # 等待5分钟(原代码逻辑) time.sleep(300) # 上传新预测文件 current_time = datetime.today().strftime("%Y-%m-%d - %Hh-%Mm-%Ss") print(f'Uploading predictions to {target_bucket_name}.') file_1_blob = upload_bucket.blob(file_1_name + '_' + current_time) file_2_blob = upload_bucket.blob(file_2_name + '_' + current_time) file_1_blob.upload_from_filename(file_1_path, content_type='text/csv') file_2_blob.upload_from_filename(file_2_path, content_type='text/csv')
排查与解决方案
1. 确保本地文件写入完成
随机丢记录最常见的原因是流水线生成文件后,文件尚未完全写入磁盘就触发上传。upload_from_filename直接读取文件路径,若文件仍在被写入(比如流水线的前序步骤未完全关闭文件句柄),会导致读取不完整。
修复方式:改用upload_from_file,通过文件对象读取,确保文件已完全写入并关闭:
# 替换原上传代码块 print(f'Uploading predictions to {target_bucket_name}.') current_time = datetime.today().strftime("%Y-%m-%d - %Hh-%Mm-%Ss") # 以二进制只读模式打开文件,确保文件已写入完成 with open(file_1_path, 'rb') as f1, open(file_2_path, 'rb') as f2: file_1_blob = upload_bucket.blob(f"{file_1_name}_{current_time}") file_1_blob.upload_from_file(f1, content_type='text/csv') file_2_blob = upload_bucket.blob(f"{file_2_name}_{current_time}") file_2_blob.upload_from_file(f2, content_type='text/csv')
2. 移除不必要的time.sleep(300)
copy_blob和blob.delete()都是同步操作,执行完成后归档和删除已彻底完成,无需等待。这个sleep不仅解决不了问题,还会浪费流水线运行时间,直接移除即可。
3. 校验文件上传完整性
通过对比本地文件和GCS文件的哈希值,确认是否是上传过程导致的数据丢失:
import hashlib def calculate_md5(file_path): hash_md5 = hashlib.md5() with open(file_path, 'rb') as f: for chunk in iter(lambda: f.read(4096), b''): hash_md5.update(chunk) return hash_md5.hexdigest() # 上传后校验 local_md5_1 = calculate_md5(file_1_path) gcs_md5_1 = file_1_blob.md5_hash if local_md5_1 != gcs_md5_1: raise Exception(f"File 1上传失败:本地MD5 {local_md5_1} != GCS MD5 {gcs_md5_1}") local_md5_2 = calculate_md5(file_2_path) gcs_md5_2 = file_2_blob.md5_hash if local_md5_2 != gcs_md5_2: raise Exception(f"File 2上传失败:本地MD5 {local_md5_2} != GCS MD5 {gcs_md5_2}")
4. 更新google-cloud-storage库
旧版本的客户端库可能存在文件上传的bug,更新到最新稳定版:
pip install --upgrade google-cloud-storage
5. 检查Vertex AI组件的文件权限
确保组件运行的服务账号对本地文件有完整的读取权限,避免因权限问题导致读取文件时截断内容。
内容的提问来源于stack exchange,提问作者AndrewJaeyoung
相关产品推荐
相关产品推荐

