You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Sagemaker Pipelines TrainingStep模型保存到指定S3路径(无唯一父文件夹)

解决SageMaker TrainingStep模型存储路径问题

要去掉TrainingStep自动生成的唯一父目录,直接将model.tar.gz存到指定S3路径,可通过以下几种方式实现:

1. 在训练脚本中手动上传模型到目标路径

绕过TrainingStep的默认输出机制,在训练完成后直接用代码将模型包上传到期望的S3位置:

import boto3
import tarfile
import os

# 打包模型文件为model.tar.gz
def package_model(model_dir, output_path):
    with tarfile.open(output_path, 'w:gz') as tar:
        tar.add(model_dir, arcname=os.path.basename(model_dir))

# 执行打包(基于SageMaker训练容器默认模型目录)
model_dir = '/opt/ml/model'
package_model(model_dir, '/tmp/model.tar.gz')

# 上传到目标S3路径
s3 = boto3.client('s3')
bucket = '{my_bucket}'
target_key = 'model/model.tar.gz'
s3.upload_file('/tmp/model.tar.gz', bucket, target_key)

这种方式下,TrainingStep的输出配置可设为临时路径(后续无需使用),核心存储逻辑由训练脚本完成。

2. 用Pipeline后处理步骤复制模型到目标路径

在TrainingStep之后添加ProcessingStep,执行S3复制操作,将默认生成的模型文件移动到指定路径:

from sagemaker.processing import ScriptProcessor
from sagemaker.workflow.steps import ProcessingStep

# 定义复制逻辑脚本
copy_script = """
import boto3
import os

s3 = boto3.client('s3')
source_bucket = os.environ['SOURCE_BUCKET']
source_key = os.environ['SOURCE_KEY']
target_bucket = os.environ['TARGET_BUCKET']
target_key = os.environ['TARGET_KEY']

# 复制模型文件到目标路径
s3.copy_object(
    Bucket=target_bucket,
    Key=target_key,
    CopySource={'Bucket': source_bucket, 'Key': source_key}
)

# 可选:删除原路径文件释放存储
s3.delete_object(Bucket=source_bucket, Key=source_key)
"""

# 创建脚本处理器
script_processor = ScriptProcessor(
    image_uri='public.ecr.aws/amazonlinux/amazonlinux:2',
    command=['python3'],
    role='your-sagemaker-role-arn',
    instance_count=1,
    instance_type='ml.t3.medium'
)

# 定义后处理步骤
copy_step = ProcessingStep(
    name='CopyModelToTarget',
    processor=script_processor,
    environment={
        'SOURCE_BUCKET': '{my_bucket}',
        'SOURCE_KEY': training_step.properties.ModelArtifacts.S3ModelArtifacts.split('/', 3)[3],
        'TARGET_BUCKET': '{my_bucket}',
        'TARGET_KEY': 'model/model.tar.gz'
    },
    code=copy_script
)

通过TrainingStep的properties.ModelArtifacts.S3ModelArtifacts获取默认模型路径,再由ProcessingStep完成复制迁移。

3. 禁用默认模型输出,完全自定义存储逻辑

定义TrainingStep时不指定model_output参数,让训练脚本全权负责模型存储:

from sagemaker.workflow.steps import TrainingStep
from sagemaker.estimator import Estimator

estimator = Estimator(
    image_uri='your-training-image-uri',
    role='your-sagemaker-role-arn',
    instance_count=1,
    instance_type='ml.m5.large',
    output_path='s3://{my_bucket}/temp-training-output'  # 临时路径,后续无需使用
)

training_step = TrainingStep(
    name='CustomTrainingStep',
    estimator=estimator,
    inputs={...}  # 你的训练输入配置
    # 不设置model_output,由训练脚本处理模型存储
)

配合第一种方法的训练脚本上传逻辑,完全绕过默认的模型输出机制。

注意事项

  • 确保训练作业的IAM角色拥有目标S3桶的PutObject、GetObject(复制场景)权限。
  • 复制完成后可删除原路径文件,避免占用额外存储资源。

内容的提问来源于stack exchange,提问作者Progress

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 02:35:05