You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Batch GPU实例无法关联ECS集群,任务卡RUNNABLE状态求助

AWS Batch GPU实例任务卡RUNNABLE状态问题排查

现有CloudFormation配置

GPULargeLaunchTemplate:
  Type: AWS::EC2::LaunchTemplate
  Properties:
    LaunchTemplateData:
      UserData:
        Fn::Base64:
          Fn::Sub: |
            MIME-Version: 1.0
            Content-Type: multipart/mixed; boundary="==BOUNDARY=="

            --==BOUNDARY==
            Content-Type: text/cloud-config; charset="us-ascii"

            runcmd:
              - yum install -y aws-cfn-bootstrap
              - echo ECS_LOGLEVEL=debug >> /etc/ecs/ecs.config
              - echo ECS_IMAGE_CLEANUP_INTERVAL=60m >> /etc/ecs/ecs.config
              - echo ECS_IMAGE_MINIMUM_CLEANUP_AGE=60m >> /etc/ecs/ecs.config
              - /opt/aws/bin/cfn-init -v --region us-west-2 --stack cool_stack --resource LaunchConfiguration
              - echo "DEVS=/dev/xvda" > /etc/sysconfig/docker-storage-setup
              - echo "VG=docker" >> /etc/sysconfig/docker-storage-setup
              - echo "DATA_SIZE=99%FREE" >> /etc/sysconfig/docker-storage-setup
              - echo "AUTO_EXTEND_POOL=yes" >> /etc/sysconfig/docker-storage-setup
              - echo "LV_ERROR_WHEN_FULL=yes" >> /etc/sysconfig/docker-storage-setup
              - echo "EXTRA_STORAGE_OPTIONS=\"--storage-opt dm.fs=ext4 --storage-opt dm.basesize=64G\"" >> /etc/sysconfig/docker-storage-setup
              - /usr/bin/docker-storage-setup
              - yum update -y
              - echo "OPTIONS=\"--default-ulimit nofile=1024000:1024000 --storage-opt dm.basesize=64G\"" >> /etc/sysconfig/docker
              - /etc/init.d/docker restart

            --==BOUNDARY==--
    LaunchTemplateName: GPULargeLaunchTemplate

GPULargeBatchComputeEnvironment:
  DependsOn:
    - ComputeRole
    - ComputeInstanceProfile
  Type: AWS::Batch::ComputeEnvironment
  Properties:
    Type: MANAGED
    ComputeResources:
      ImageId: ami-GPU-optimized-AMI-ID
      AllocationStrategy: BEST_FIT_PROGRESSIVE
      LaunchTemplate:
        LaunchTemplateId:
          Ref: GPULargeLaunchTemplate
        Version:
          Fn::GetAtt:
            - GPULargeLaunchTemplate
            - LatestVersionNumber
      InstanceRole:
        Ref: ComputeInstanceProfile
      InstanceTypes:
        - g4dn.xlarge
      MaxvCpus: 768
      MinvCpus: 1
      SecurityGroupIds:
        - Fn::GetAtt:
            - ComputeSecurityGroup
            - GroupId
      Subnets:
        - Ref: ComputePrivateSubnetA
      Type: EC2
      UpdateToLatestImageVersion: True

MyGPUBatchJobQueue:
  Type: AWS::Batch::JobQueue
  Properties:
    ComputeEnvironmentOrder:
      - ComputeEnvironment:
          Ref: GPULargeBatchComputeEnvironment
        Order: 1
    Priority: 5
    JobQueueName: MyGPUBatchJobQueue
    State: ENABLED

MyGPUJobDefinition:
  Type: AWS::Batch::JobDefinition
  Properties:
    Type: container
    ContainerProperties:
      Command:
        - "/opt/bin/python3"
        - "/opt/bin/start.py"
        - "--retry_count"
        - "Ref::batchRetryCount"
        - "--retry_limit"
        - "Ref::batchRetryLimit"
      Environment:
        - Name: "Region"
          Value: "us-west-2"
        - Name: "LANG"
          Value: "en_US.UTF-8"
      Image:
        Fn::Sub: "cool_1234_abc.dkr.ecr.us-west-2.amazonaws.com/my-image"
      JobRoleArn:
        Fn::Sub: "arn:aws:iam::cool_1234_abc:role/ComputeRole"
      Memory: 16000
      Vcpus: 1
      ResourceRequirements:
        - Type: GPU
          Value: '1'
    JobDefinitionName: MyGPUJobDefinition
    Timeout:
      AttemptDurationSeconds: 500

已完成的排查步骤

  • 替换为普通CPU实例后任务可正常运行,确认问题出在GPU实例配置环节
  • 使用GPU优化AMI后问题仍未解决
  • 对比CPU与GPU任务的describe-jobs结果,GPU任务缺失containerInstanceArn和taskArn字段
  • ASG中存在GPU实例,但对应ECS集群未关联任何容器实例(CPU实例则正常关联)

解决思路

1. 检查ECS Agent集群注册配置

AWS Batch托管计算环境会自动创建专属ECS集群,GPU实例需正确注册到该集群:

  • 登录GPU实例,查看/etc/ecs/ecs.config文件,确认是否包含ECS_CLUSTER参数,且值与Batch计算环境对应的ECS集群名称一致(可通过Batch控制台的计算环境详情获取集群名)
  • 若UserData中未设置该参数,需添加echo ECS_CLUSTER=<集群名称> >> /etc/ecs/ecs.config到runcmd列表

2. 验证ECS Agent运行状态与日志

登录GPU实例执行以下命令排查:

  • sudo systemctl status ecs:检查ECS Agent是否正常运行
  • sudo cat /var/log/ecs/ecs-agent.log:查看Agent日志,定位注册失败、权限不足或资源初始化错误信息
  • 确认Docker重启后ECS Agent是否自动重启,若未重启手动执行sudo systemctl restart ecs

3. 确认GPU驱动与容器运行时兼容性

GPU优化AMI需确保驱动与nvidia-docker组件适配:

  • 执行nvidia-smi验证GPU驱动是否正常加载
  • 执行docker run --rm nvidia/cuda:11.0-base nvidia-smi验证容器能否访问GPU资源
  • 若AMI未预装nvidia-docker,在UserData中添加安装步骤:
    - amazon-linux-extras install -y nvidia-docker
    - systemctl restart docker
    

4. 检查实例权限与网络连通性

  • 确认ComputeInstanceProfile关联的IAM角色包含AmazonEC2ContainerServiceforEC2Role权限,该权限允许ECS Agent向ECS服务注册实例
  • 检查GPU实例所在私有子网是否配置NAT网关,确保实例能访问AWS公网服务(ECS Agent需连接ECS服务端点)
  • 验证安全组是否允许实例出站HTTPS(443端口)访问

5. 核对Batch计算环境与任务定义的GPU配置

  • 确认g4dn.xlarge实例的GPU资源数量与任务定义中ResourceRequirements的Value: '1'匹配
  • 检查计算环境的UpdateToLatestImageVersion设置是否正确加载最新GPU优化AMI

内容的提问来源于stack exchange,提问作者Fisher Coder

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 01:15:57