如何指定SageMaker TensorFlow Estimator训练源码的S3上传路径以匹配output_path?
要让训练源码上传到output_path对应的存储桶和前缀下,关键是显式设置Estimator的code_location参数,而不是依赖默认行为。
问题原因
默认情况下,当你只指定output_path时,SageMaker会将训练源码上传到存储桶根目录下的/<training_job_name>/source路径,不会自动继承output_path里的前缀。这就是你看到源码路径和模型输出路径层级不一致的核心原因。
解决方案
在初始化TensorFlow Estimator时,添加code_location参数,将其设置为output_path对应的S3路径(即s3://<bucket>/<prefix>/)。这样SageMaker会把源码上传到code_location下的/<training_job_name>/source目录,和模型输出的/<training_job_name>/output保持同层级,都落在你指定的前缀下。
代码示例
import sagemaker from sagemaker.tensorflow import TensorFlow # 定义你的模型输出路径 output_path = "s3://<bucket>/<prefix>/" # 初始化Estimator时显式指定code_location estimator = TensorFlow( entry_point="train.py", role=sagemaker.get_execution_role(), instance_count=1, instance_type="ml.p3.2xlarge", framework_version="2.12", output_path=output_path, # 关键配置:让源码路径继承output_path的桶和前缀 code_location=output_path ) # 启动训练作业 estimator.fit()
完成设置后,源码会被上传到s3://<bucket>/<prefix>/<training_job_name>/source,模型产物则在s3://<bucket>/<prefix>/<training_job_name>/output,完全符合你的预期路径结构。
灵活扩展
如果需要给源码路径单独添加子前缀(比如在指定前缀下再加source_code目录),可以这样设置:
code_location = f"{output_path}source_code/"
此时源码会被上传到s3://<bucket>/<prefix>/source_code/<training_job_name>/source,满足更细分的存储需求。
内容的提问来源于stack exchange,提问作者Austin

