You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SageMaker Pipeline端点部署失败:CannotStartContainerError求助

SageMaker Pipeline部署端点失败,但直接SDK部署成功的问题

我搭建了一个SageMaker Pipeline,包含训练、创建模型、Lambda部署端点三个步骤,核心代码如下:

# 1. 模型训练步骤
estimator = TensorFlow(
    entry_point="train.py",
    source_dir=src_dir,
    role=role,
    instance_count=1,
    instance_type="ml.m4.4xlarge",
    framework_version="2.1",
    py_version="py3",
    base_job_name="quantitative-scores-training",
    output_path=s3_training_output_file,
    code_location=f"{base_dir}/code/"
)

training_inputs = {
    'train': TrainingInput(
        s3_data=s3_training_data_input_file,
        content_type='text/csv',
        input_mode='FastFile'
    )
}

training_step = TrainingStep(
    name='Train',
    estimator=estimator,
    inputs=training_inputs,
)

# 2. 创建模型步骤
model = Model(
    entry_point='inference.py',
    source_dir=src_dir,
    model_data=training_step.properties.ModelArtifacts.S3ModelArtifacts,
    role=role,
    sagemaker_session=sagemaker_session,
    image_uri=estimator.training_image_uri(),
)

create_model_step = ModelStep(
    name="ModelStep",
    step_args=model.create(
        instance_type='ml.m4.4xlarge'
    ),
)

# 3. 部署模型到端点步骤
deploy_model_lambda_function = Lambda(
    function_name="sagemaker-deploy-quant-score",
    execution_role_arn=create_sagemaker_lambda_role("deploy-model-lambda-role"),
    script="/home/ec2-user/SageMaker/my_path/src/util/deploy_model_lambda.py",
    handler="deploy_model_lambda.lambda_handler",
)

deploy_model_step = LambdaStep(
    name="DeployModelStep",
    lambda_func=deploy_model_lambda_function,
    inputs={
        "model_name": create_model_step.properties.ModelName,
        "endpoint_config_name": "quantitative-scoring-pipeline-config",
        "endpoint_name": endpoint_name,
        "endpoint_instance_type": "ml.m4.xlarge",
    },
)

# 构建并启动Pipeline
pipe = Pipeline(
    name="QuantitativeScoringPipeline",
    steps=[
        training_step,
        create_model_step,
        deploy_model_step
    ],
    parameters=[
        # 省略参数定义
        s3_training_data_input_file,
        s3_training_output_file,
        endpoint_name
    ],
)
pipe.upsert(role_arn=role)
execution = pipe.start()

运行后Lambda执行成功,但端点创建失败,报错:

CannotStartContainerError. Please ensure the model container for variant AllTraffic starts correctly when invoked with 'docker run serve'

容器从未启动,CloudWatch中也没有相关日志。

奇怪的是,我用以下SageMaker SDK代码直接部署同一个模型S3 URI却完全正常:

model = TensorFlowModel(
    entry_point='inference.py',
    source_dir='src',
    model_data="s3://sagemaker-eu-west-1-558091818291/tensorflow-training-2024-04-25-12-11-21-401/pipelines-dgstz6rrp8u9-ModelStep-RepackMode-P5O95TSntC/output/model.tar.gz",
    role=role,
    framework_version="2.1",
)
predictor = model.deploy(instance_type='ml.m4.xlarge', initial_instance_count=1, endpoint_name=endpoint_name)

不过这种方式会生成新的模型压缩包,我对比过两个压缩包的内容,推理代码和模型数据完全一致。目前能想到的差异点是Pipeline中用了训练镜像URI,而直接部署用的是框架版本,但不知道该怎么解决这个问题。

内容的提问来源于stack exchange,提问作者Tom111989

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 05:53:24