如何将Sagemaker Pipelines TrainingStep模型保存到指定S3路径(无唯一父文件夹)
解决SageMaker TrainingStep模型存储路径问题
要去掉TrainingStep自动生成的唯一父目录,直接将model.tar.gz存到指定S3路径,可通过以下几种方式实现:
1. 在训练脚本中手动上传模型到目标路径
绕过TrainingStep的默认输出机制,在训练完成后直接用代码将模型包上传到期望的S3位置:
import boto3 import tarfile import os # 打包模型文件为model.tar.gz def package_model(model_dir, output_path): with tarfile.open(output_path, 'w:gz') as tar: tar.add(model_dir, arcname=os.path.basename(model_dir)) # 执行打包(基于SageMaker训练容器默认模型目录) model_dir = '/opt/ml/model' package_model(model_dir, '/tmp/model.tar.gz') # 上传到目标S3路径 s3 = boto3.client('s3') bucket = '{my_bucket}' target_key = 'model/model.tar.gz' s3.upload_file('/tmp/model.tar.gz', bucket, target_key)
这种方式下,TrainingStep的输出配置可设为临时路径(后续无需使用),核心存储逻辑由训练脚本完成。
2. 用Pipeline后处理步骤复制模型到目标路径
在TrainingStep之后添加ProcessingStep,执行S3复制操作,将默认生成的模型文件移动到指定路径:
from sagemaker.processing import ScriptProcessor from sagemaker.workflow.steps import ProcessingStep # 定义复制逻辑脚本 copy_script = """ import boto3 import os s3 = boto3.client('s3') source_bucket = os.environ['SOURCE_BUCKET'] source_key = os.environ['SOURCE_KEY'] target_bucket = os.environ['TARGET_BUCKET'] target_key = os.environ['TARGET_KEY'] # 复制模型文件到目标路径 s3.copy_object( Bucket=target_bucket, Key=target_key, CopySource={'Bucket': source_bucket, 'Key': source_key} ) # 可选:删除原路径文件释放存储 s3.delete_object(Bucket=source_bucket, Key=source_key) """ # 创建脚本处理器 script_processor = ScriptProcessor( image_uri='public.ecr.aws/amazonlinux/amazonlinux:2', command=['python3'], role='your-sagemaker-role-arn', instance_count=1, instance_type='ml.t3.medium' ) # 定义后处理步骤 copy_step = ProcessingStep( name='CopyModelToTarget', processor=script_processor, environment={ 'SOURCE_BUCKET': '{my_bucket}', 'SOURCE_KEY': training_step.properties.ModelArtifacts.S3ModelArtifacts.split('/', 3)[3], 'TARGET_BUCKET': '{my_bucket}', 'TARGET_KEY': 'model/model.tar.gz' }, code=copy_script )
通过TrainingStep的properties.ModelArtifacts.S3ModelArtifacts获取默认模型路径,再由ProcessingStep完成复制迁移。
3. 禁用默认模型输出,完全自定义存储逻辑
定义TrainingStep时不指定model_output参数,让训练脚本全权负责模型存储:
from sagemaker.workflow.steps import TrainingStep from sagemaker.estimator import Estimator estimator = Estimator( image_uri='your-training-image-uri', role='your-sagemaker-role-arn', instance_count=1, instance_type='ml.m5.large', output_path='s3://{my_bucket}/temp-training-output' # 临时路径,后续无需使用 ) training_step = TrainingStep( name='CustomTrainingStep', estimator=estimator, inputs={...} # 你的训练输入配置 # 不设置model_output,由训练脚本处理模型存储 )
配合第一种方法的训练脚本上传逻辑,完全绕过默认的模型输出机制。
注意事项
- 确保训练作业的IAM角色拥有目标S3桶的
PutObject、GetObject(复制场景)权限。 - 复制完成后可删除原路径文件,避免占用额外存储资源。
内容的提问来源于stack exchange,提问作者Progress
相关产品推荐
相关产品推荐

