解决SageMaker+FSx for Lustre训练时工件上传失败问题
修复SageMaker训练任务Artifact上传失败问题
问题描述
我在实现基于Amazon SageMaker和FSx for Lustre的迁移学习方案时,训练任务报错失败:
UnexpectedStatusException: Error for Training job tensorflow-training-2023-04-14-XX-XX-XX-XXX: Failed. Reason: ClientError: Artifact upload failed:Please ensure that the subnet's route table has a route to an S3 VPC endpoint or a NAT device, and both the security groups and the subnet's network ACL allow uploading data to all output URIs
当前环境信息
- 已成功配置FSx for Lustre文件系统
- 文件系统已关联到目标S3存储桶的指定文件夹
- 使用的是启用ACL的S3存储桶
- 子网路由表条目:
Destination Target 172.31.0.0/16 local 0.0.0.0/0 igw-XXXXXXXX
疑问点
- 什么是S3 VPC端点?如何检查它的运行状态及路由表关联情况?
- 如何确保安全组和子网网络ACL允许向输出URI上传数据?
关联代码
from sagemaker.inputs import FileSystemInput # Specify file system id. file_system_id = "fs-XXXXXXXXXXXXXXXXX" #FSx_SM_Input # Specify directory path associated with the file system. You need to provide normalized and absolute path here. file_system_directory_path = "/YYYYYYY/XXXX" # Specify the access mode of the mount of the directory associated with the file system. # Directory can be mounted either in 'ro'(read-only) or 'rw' (read-write). file_system_access_mode = "rw" # Specify your file system type, "EFS" or "FSxLustre". file_system_type = "FSxLustre" # Give Amazon SageMaker Training Jobs Access to FileSystem Resources in Your Amazon VPC. security_groups_ids = ["sg-XXXXXXX"] subnets = ["subnet-XXXXXXXX"] fs_train_input = FileSystemInput( file_system_id=file_system_id, file_system_type=file_system_type, directory_path=file_system_directory_path, file_system_access_mode=file_system_access_mode, ) import sagemaker from sagemaker import get_execution_role from sagemaker.tensorflow import TensorFlow hyperparameters = {"input_mode": "File", #"FastFile", "Pipe", "File" -- FOR FSx "File" ONLY!!! "shards_on_input": 4 # 4, 120 -- Shards in S3 } train_input = sagemaker.inputs.TrainingInput("s3://mnist-tdrecords/train/{}".format(hyperparameters["shards_on_input"]), input_mode = hyperparameters["input_mode"], distribution = 'FullyReplicated' #'ShardedByS3Key', 'FullyReplicated' ) tf_estimator = TensorFlow(entry_point = "AWS_DataPipping_TFMirroredStrategy.py", source_dir = "./", framework_version = "2.3", py_version = "py37", instance_type = "ml.p3.2xlarge", # "ml.p3.2xlarge", "ml.p3.8xlarge", "ml.p3.16xlarge" instance_count = 1, role = sagemaker.get_execution_role(), subnets=subnets, security_group_ids=security_groups_ids, hyperparameters = hyperparameters, output_path = f"s3://mnist-tdrecords/output", input_mode = hyperparameters["input_mode"], # "File", "Pipe", "FastFile") ) s3_data_channels = {"train": fs_train_input} #s3_data_channels = {"train": "s3://mnist-tdrecords/train/{}".format(hyperparameters["shards_on_input"])} #"validation": f"s3://{bucket_name}/data/validation",} tf_estimator.fit(s3_data_channels)
解决方案与故障排查步骤
一、网络连通性修复(路由表与NAT/S3 VPC端点)
错误核心是训练实例无法访问S3上传结果,针对你的路由表配置:
- NAT网关检查:
- 如果你的子网是私有子网,仅配置互联网网关(IGW)无法让私有实例访问公网S3,必须部署NAT网关,并将路由表中
0.0.0.0/0的目标修改为NAT网关ID,而非IGW。 - 如果是公有子网,需确保训练实例自动分配了公网IP,否则无法通过IGW访问S3。
- 如果你的子网是私有子网,仅配置互联网网关(IGW)无法让私有实例访问公网S3,必须部署NAT网关,并将路由表中
- S3 VPC端点配置:
- S3网关型VPC端点是VPC内实例无需公网直接访问S3的免费方案,优先级高于公网路由。
- 检查与创建步骤:
- 进入VPC控制台 → 「端点」页面,搜索是否存在类型为
Gateway、服务名为com.amazonaws.<你的区域>.s3的端点 - 确认该端点已关联到训练子网所在的路由表(关联后路由表会自动添加S3前缀列表
pl-xxxxxx指向端点ID的条目) - 若未创建,直接创建S3网关端点,关联目标路由表即可
- 进入VPC控制台 → 「端点」页面,搜索是否存在类型为
二、安全组与网络ACL验证
安全组(实例层面)
- 出站规则:添加允许TCP 443端口(HTTPS)到
0.0.0.0/0或S3前缀列表的规则,确保实例能发起S3请求。 - 入站规则:无需额外配置,保持默认拒绝即可(训练实例不需要外部主动访问)。
网络ACL(子网层面)
- 出站规则:确保允许TCP 443端口流量到
0.0.0.0/0或S3前缀列表,规则编号要低于默认的拒绝所有规则(默认ACL允许所有出站,若有自定义规则需检查优先级)。 - 入站规则:允许来自S3的TCP 443端口响应流量,若使用S3 VPC端点,只需允许VPC内部CIDR的流量即可。
三、S3存储桶权限补充
由于你的存储桶启用了ACL,需确保SageMaker执行角色拥有:
s3:PutObject权限,目标资源为arn:aws:s3:::mnist-tdrecords/output/*- 同时检查存储桶ACL是否允许该角色所属AWS账号拥有写入权限,或直接在存储桶策略中明确授权该角色的写入权限。
四、代码配置确认
- 确认指定的
subnets如果是私有子网,已配置NAT网关或S3 VPC端点 - 确认
security_group_ids对应的安全组出站规则无限制(或允许S3访问)
内容的提问来源于stack exchange,提问作者sebtac
相关产品推荐
相关产品推荐

