You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python3将字节数据上传至GCP并实现S3到GCP文件迁移

直接通过字节数据将S3文件上传至GCP存储桶

是的,完全可以直接使用从S3获取的字节数据上传到GCP Cloud Storage,不需要将文件保存到本地磁盘。你的代码目前存在一处关键调用错误,以下是修正后的完整实现方案:

问题修正与核心逻辑

原代码中blob(source_file, if_generation_match=generation_match_precondition)是错误的调用方式,GCP的Blob对象没有这种直接调用的方法。针对内存中的字节数据,我们可以使用upload_from_string()方法直接传入字节串,或者将字节包装成BytesIO对象后使用upload_from_file()方法。

完整修正代码

import os
from io import BytesIO
import boto3
from google.cloud import storage

# 需提前定义的变量:bucket_name, folder, gcp_bucket_name, role_arn, role_arn_name

def build_session(role_arn, role_session_name):
    # 补充STS角色凭证获取逻辑
    sts_client = boto3.client('sts')
    assumed_role_object = sts_client.assume_role(
        RoleArn=role_arn,
        RoleSessionName=role_session_name
    )
    credentials = assumed_role_object['Credentials']
    return boto3.Session(
        aws_access_key_id=credentials['AccessKeyId'],
        aws_secret_access_key=credentials['SecretAccessKey'],
        aws_session_token=credentials['SessionToken']
    )

def parse_type(object_key):
    # 补充自定义文件类型解析逻辑
    if object_key.endswith('.csv'):
        return 'csv/'
    elif object_key.endswith('.json'):
        return 'json/'
    return ''

def download_from_s3():
    s3_client = build_session(role_arn, role_arn_name).client('s3')
    files = s3_client.list_objects_v2(
        Bucket=bucket_name,
        Prefix=folder
    )

    output = []
    for content in files.get('Contents', []):
        object_key = content['Key']
        report_type = parse_type(object_key)
        if report_type == "":
            continue
                    
        file = s3_client.get_object(
            Bucket=bucket_name,
            Key=object_key
        )

        object_body = file['Body'].read()
        output.append({
            "Key": os.path.basename(object_key),
            "Type": report_type,
            "File": object_body,
        })

    return output

def upload_blob(bucket_name, source_bytes, destination_blob_name):
    """使用字节数据上传文件至GCP存储桶"""
    storage_client = storage.Client()
    bucket = storage_client.bucket(bucket_name)
    blob = bucket.blob(destination_blob_name)

    # 生成匹配预条件,避免覆盖已有数据
    generation_match_precondition = 0

    # 直接上传字节数据
    blob.upload_from_string(
        source_bytes,
        if_generation_match=generation_match_precondition
    )

    print(f"字节数据已上传至 {destination_blob_name}")

def send_from_s3_to_gcp(s3_files):
    for content in s3_files:
        upload_blob(
            bucket_name=gcp_bucket_name,
            source_bytes=content["File"],
            destination_blob_name=content["Type"] + content["Key"],
        )

# 执行流程
s3_files = download_from_s3()
send_from_s3_to_gcp(s3_files)

关键细节说明

  • upload_from_string():这是处理内存字节数据最直接的方法,支持直接传入bytes类型数据完成上传。
  • 替代方案:如果需要模拟文件对象上传,可将字节数据包装为BytesIO对象,调用blob.upload_from_file(BytesIO(source_bytes)),效果一致。
  • 预条件设置:保留if_generation_match=0确保仅当目标对象不存在时才上传,避免意外覆盖,可根据需求调整该值。

内容的提问来源于stack exchange,提问作者Ming

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 19:52:22