如何用Python3将字节数据上传至GCP并实现S3到GCP文件迁移
直接通过字节数据将S3文件上传至GCP存储桶
是的,完全可以直接使用从S3获取的字节数据上传到GCP Cloud Storage,不需要将文件保存到本地磁盘。你的代码目前存在一处关键调用错误,以下是修正后的完整实现方案:
问题修正与核心逻辑
原代码中blob(source_file, if_generation_match=generation_match_precondition)是错误的调用方式,GCP的Blob对象没有这种直接调用的方法。针对内存中的字节数据,我们可以使用upload_from_string()方法直接传入字节串,或者将字节包装成BytesIO对象后使用upload_from_file()方法。
完整修正代码
import os from io import BytesIO import boto3 from google.cloud import storage # 需提前定义的变量:bucket_name, folder, gcp_bucket_name, role_arn, role_arn_name def build_session(role_arn, role_session_name): # 补充STS角色凭证获取逻辑 sts_client = boto3.client('sts') assumed_role_object = sts_client.assume_role( RoleArn=role_arn, RoleSessionName=role_session_name ) credentials = assumed_role_object['Credentials'] return boto3.Session( aws_access_key_id=credentials['AccessKeyId'], aws_secret_access_key=credentials['SecretAccessKey'], aws_session_token=credentials['SessionToken'] ) def parse_type(object_key): # 补充自定义文件类型解析逻辑 if object_key.endswith('.csv'): return 'csv/' elif object_key.endswith('.json'): return 'json/' return '' def download_from_s3(): s3_client = build_session(role_arn, role_arn_name).client('s3') files = s3_client.list_objects_v2( Bucket=bucket_name, Prefix=folder ) output = [] for content in files.get('Contents', []): object_key = content['Key'] report_type = parse_type(object_key) if report_type == "": continue file = s3_client.get_object( Bucket=bucket_name, Key=object_key ) object_body = file['Body'].read() output.append({ "Key": os.path.basename(object_key), "Type": report_type, "File": object_body, }) return output def upload_blob(bucket_name, source_bytes, destination_blob_name): """使用字节数据上传文件至GCP存储桶""" storage_client = storage.Client() bucket = storage_client.bucket(bucket_name) blob = bucket.blob(destination_blob_name) # 生成匹配预条件,避免覆盖已有数据 generation_match_precondition = 0 # 直接上传字节数据 blob.upload_from_string( source_bytes, if_generation_match=generation_match_precondition ) print(f"字节数据已上传至 {destination_blob_name}") def send_from_s3_to_gcp(s3_files): for content in s3_files: upload_blob( bucket_name=gcp_bucket_name, source_bytes=content["File"], destination_blob_name=content["Type"] + content["Key"], ) # 执行流程 s3_files = download_from_s3() send_from_s3_to_gcp(s3_files)
关键细节说明
upload_from_string():这是处理内存字节数据最直接的方法,支持直接传入bytes类型数据完成上传。- 替代方案:如果需要模拟文件对象上传,可将字节数据包装为
BytesIO对象,调用blob.upload_from_file(BytesIO(source_bytes)),效果一致。 - 预条件设置:保留
if_generation_match=0确保仅当目标对象不存在时才上传,避免意外覆盖,可根据需求调整该值。
内容的提问来源于stack exchange,提问作者Ming
相关产品推荐
相关产品推荐

