You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练流程在下载输入数据阶段失败求助

问题:Amazon SageMaker训练任务在输入数据下载阶段异常停止

自定义Docker镜像

FROM 683313688378.dkr.ecr.us-east-1.amazonaws.com/sagemaker-scikit-learn:0.20.0-cpu-py3

RUN pip install -U spacy boto3

RUN python -m spacy download en_core_web_sm

训练任务创建代码

import boto3
import sagemaker
from sagemaker import get_execution_role

container_uri = "..."
role = get_execution_role()

model = sagemaker.estimator.Estimator(
        image_uri=container_uri,
        role=role,
        instance_count=1,
        instance_type="ml.m4.xlarge",
        volume_size=1,
        entry_point="train.py",
        source_dir="spacy-train-custom",
        dependencies=["spacy-train-custom/configs"],
        output_path=f"s3://test-bucket/output/"
)

model.fit()

训练任务执行日志

Using role ...
2022-10-06 00:21:15 Starting - Starting the training job...
2022-10-06 00:21:40 Starting - Preparing the instances for trainingProfilerReport-1665015675: InProgress
.........
2022-10-06 00:23:09 Downloading - Downloading input data
2022-10-06 00:23:09 Stopping - Stopping the training job
2022-10-06 00:23:09 Stopped - Training job stopped
ProfilerReport-1665015675: Stopping
..
Job ended with status 'Stopped' rather than 'Completed'. This could mean the job timed out or stopped early for some other reason: Consider checking whether it completed as you expect.

训练脚本代码

import os
from sagemaker_training import environment
from spacy.cli.train import train

print("start training")

env = environment.Environment()
config_path = env.hyperparameters.pop("config")

overrides = {
    "paths.train": os.path.join(env.channel_input_dirs["train"], "train.spacy"),
}
overrides.update(env.hyperparameters)

use_gpu: int = 0 if env.num_gpus > 0 else -1

train(
    config_path,
    output_path=env.model_dir,
    use_gpu=use_gpu,
    overrides=overrides,
)

排查方向

  • 输入数据配置问题:
    即使添加了输入配置,需确认:

    • 输入通道指定的S3路径是否存在,无拼写错误;
    • 训练角色是否具备访问该S3 bucket的s3:GetObject等权限;
    • 若使用model.fit(inputs=...),需确保输入通道的key(比如"train")和脚本中env.channel_input_dirs["train"]一致。
  • Docker镜像依赖缺失:
    基础镜像sagemaker-scikit-learn:0.20.0-cpu-py3可能未预装sagemaker-training库,而训练脚本中直接导入了该库,会导致脚本启动失败,触发任务停止。需在Dockerfile中添加RUN pip install sagemaker-training。

  • 资源配置不足:
    当前设置的volume_size=1(GB)过小,若输入数据体积较大,会导致磁盘空间不足,下载失败后任务停止。建议调整为更大值,比如volume_size=10。

  • 镜像兼容性问题:
    基础镜像版本较旧(scikit-learn 0.20.0),可能与当前使用的SageMaker SDK版本不兼容,导致训练流程出现异常。可尝试升级基础镜像到较新版本,比如sagemaker-scikit-learn:1.2-1-cpu-py3。

  • 日志细节排查:
    登录SageMaker控制台,进入该训练任务详情页,查看CloudWatch日志组中的具体报错信息,比如S3访问错误、磁盘空间告警、脚本启动异常等,这些是定位问题的关键。

内容的提问来源于stack exchange,提问作者arielnmz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 12:15:44