You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在AWS Batch作业定义JSON中指定输入输出S3卷路径?

AWS Batch作业定义配置指南(EC2/Fargate后端)

问题背景

我需要在AWS Batch上用EC2或Fargate后端运行VEBA预处理作业,本地Docker运行命令如下:

# Directories
LOCAL_WORKING_DIRECTORY=$(pwd)
LOCAL_OUTPUT_PARENT_DIRECTORY=../
LOCAL_OUTPUT_PARENT_DIRECTORY=$(realpath -m ${LOCAL_OUTPUT_PARENT_DIRECTORY})

CONTAINER_INPUT_DIRECTORY=/volumes/input/
CONTAINER_OUTPUT_DIRECTORY=/volumes/output/

# Parameters
ID=S1
R1=Fastq/${ID}_1.fastq.gz
R2=Fastq/${ID}_2.fastq.gz
NAME=VEBA-preprocess__${ID}
RELATIVE_OUTPUT_DIRECTORY=veba_output/preprocess/

# Command
CMD="preprocess.py -1 ${CONTAINER_INPUT_DIRECTORY}/${R1} -2 ${CONTAINER_INPUT_DIRECTORY}/${R2} -n ${ID} -o ${CONTAINER_OUTPUT_DIRECTORY}/${RELATIVE_OUTPUT_DIRECTORY}"

# Docker
DOCKER_IMAGE="jolespin/veba_preprocess:1.1.2"
docker run \
    --name ${NAME} \
    --rm \
    --volume ${LOCAL_WORKING_DIRECTORY}:${CONTAINER_INPUT_DIRECTORY} \
    --volume ${LOCAL_OUTPUT_PARENT_DIRECTORY}:${CONTAINER_OUTPUT_DIRECTORY} \
    ${DOCKER_IMAGE} \
    -c "${CMD}"

当前输入文件存于S3路径s3://path/to/input/(包含35_R1.fq.gz和35_R2.fq.gz),需将输出存储到s3://path/to/output/,但不清楚如何在AWS Batch作业定义JSON中配置S3路径的卷挂载,现有不完整的作业定义如下:

{
  "jobDefinitionName": "preprocess__35",
  "type": "container",
  "containerProperties": {
    "image": "jolespin/veba_preprocess:1.1.2",
    "vcpus": 4,
    "memory": 16000,
    "command": [
      "preprocess.py",
      "-1",
      "/volumes/input/35_R1.fq.gz",
      "-2",
      "/volumes/input/35_R2.fq.gz",
      "-n",
      "35",
      "-o",
      "/volumes/output/veba_output/preprocess",
      "-p",
      "4"
    ],
    "mountPoints": [
      {

      }
    ],
    "volumes": [
    
      
    ]
  }
}

解决方案

AWS Batch本身不支持直接将S3作为卷挂载到容器,需根据后端类型(EC2/Fargate)选择合适的方式处理S3输入输出,以下是两种后端的完整配置:

一、EC2后端配置

EC2后端可选两种方案:利用实例本地存储+S3同步,或使用EFS持久化存储。

方案1:实例本地存储+S3同步(轻量快速)

直接在作业命令中完成S3输入下载、预处理、输出上传的流程,无需额外存储服务:

{
  "jobDefinitionName": "preprocess__35",
  "type": "container",
  "platformCapabilities": ["EC2"],
  "containerProperties": {
    "image": "jolespin/veba_preprocess:1.1.2",
    "vcpus": 4,
    "memory": 16000,
    "command": [
      "/bin/bash",
      "-c",
      "mkdir -p /volumes/input /volumes/output && aws s3 sync s3://path/to/input/ /volumes/input/ && preprocess.py -1 /volumes/input/35_R1.fq.gz -2 /volumes/input/35_R2.fq.gz -n 35 -o /volumes/output/veba_output/preprocess -p 4 && aws s3 sync /volumes/output/ s3://path/to/output/"
    ],
    "mountPoints": [],
    "volumes": [],
    "jobRoleArn": "arn:aws:iam::你的账号ID:role/BatchJobS3AccessRole"
  },
  "retryStrategy": {
    "attempts": 1
  }
}
  • 替换jobRoleArn为你的IAM角色ARN,该角色需拥有s3:GetObject、s3:PutObject权限,允许访问指定S3路径。
  • 用/bin/bash -c串联命令:创建容器内目录、同步S3输入到容器、执行预处理、同步输出回S3。

方案2:挂载EC2实例本地卷

如果需要使用实例本地存储作为临时卷,可配置如下:

{
  "jobDefinitionName": "preprocess__35",
  "type": "container",
  "platformCapabilities": ["EC2"],
  "containerProperties": {
    "image": "jolespin/veba_preprocess:1.1.2",
    "vcpus": 4,
    "memory": 16000,
    "command": [
      "/bin/bash",
      "-c",
      "aws s3 sync s3://path/to/input/ /volumes/input/ && preprocess.py -1 /volumes/input/35_R1.fq.gz -2 /volumes/input/35_R2.fq.gz -n 35 -o /volumes/output/veba_output/preprocess -p 4 && aws s3 sync /volumes/output/ s3://path/to/output/"
    ],
    "mountPoints": [
      {
        "sourceVolume": "input-volume",
        "containerPath": "/volumes/input",
        "readOnly": false
      },
      {
        "sourceVolume": "output-volume",
        "containerPath": "/volumes/output",
        "readOnly": false
      }
    ],
    "volumes": [
      {
        "host": {
          "sourcePath": "/tmp/input"
        },
        "name": "input-volume"
      },
      {
        "host": {
          "sourcePath": "/tmp/output"
        },
        "name": "output-volume"
      }
    ],
    "jobRoleArn": "arn:aws:iam::你的账号ID:role/BatchJobS3AccessRole"
  }
}
  • 这里使用EC2实例的/tmp目录作为卷,作业结束后会自动清理,适合临时存储。

二、Fargate后端配置

Fargate不支持实例本地卷,推荐使用S3同步方案,部分区域也支持挂载EFS。

方案1:S3同步(通用推荐)

配置方式类似EC2,但需添加Fargate专属的执行角色和网络配置:

{
  "jobDefinitionName": "preprocess__35",
  "type": "container",
  "platformCapabilities": ["FARGATE"],
  "containerProperties": {
    "image": "jolespin/veba_preprocess:1.1.2",
    "vcpus": 4,
    "memory": 16000,
    "command": [
      "/bin/bash",
      "-c",
      "mkdir -p /volumes/input /volumes/output && aws s3 sync s3://path/to/input/ /volumes/input/ && preprocess.py -1 /volumes/input/35_R1.fq.gz -2 /volumes/input/35_R2.fq.gz -n 35 -o /volumes/output/veba_output/preprocess -p 4 && aws s3 sync /volumes/output/ s3://path/to/output/"
    ],
    "mountPoints": [],
    "volumes": [],
    "executionRoleArn": "arn:aws:iam::你的账号ID:role/BatchFargateExecutionRole",
    "jobRoleArn": "arn:aws:iam::你的账号ID:role/BatchJobS3AccessRole",
    "networkConfiguration": {
      "awsvpcConfiguration": {
        "subnets": ["subnet-你的子网ID"],
        "securityGroups": ["sg-你的安全组ID"],
        "assignPublicIp": "ENABLED"
      }
    }
  },
  "retryStrategy": {
    "attempts": 1
  }
}
  • executionRoleArn是Fargate执行角色,需拥有拉取镜像、写入CloudWatch日志的权限。
  • networkConfiguration需配置你的VPC子网和安全组,assignPublicIp设为ENABLED让任务能访问S3(或使用VPC端点)。

方案2:挂载EFS(限支持区域)

若你的AWS区域支持Fargate挂载EFS,可先将S3内容同步到EFS,再挂载到容器:

{
  "jobDefinitionName": "preprocess__35",
  "type": "container",
  "platformCapabilities": ["FARGATE"],
  "containerProperties": {
    "image": "jolespin/veba_preprocess:1.1.2",
    "vcpus": 4,
    "memory": 16000,
    "command": [
      "/bin/bash",
      "-c",
      "aws s3 sync s3://path/to/input/ /volumes/input/ && preprocess.py -1 /volumes/input/35_R1.fq.gz -2 /volumes/input/35_R2.fq.gz -n 35 -o /volumes/output/veba_output/preprocess -p 4 && aws s3 sync /volumes/output/ s3://path/to/output/"
    ],
    "mountPoints": [
      {
        "sourceVolume": "efs-input",
        "containerPath": "/volumes/input",
        "readOnly": false
      },
      {
        "sourceVolume": "efs-output",
        "containerPath": "/volumes/output",
        "readOnly": false
      }
    ],
    "volumes": [
      {
        "efsVolumeConfiguration": {
          "fileSystemId": "fs-你的EFS ID",
          "rootDirectory": "/input",
          "transitEncryption": "ENABLED"
        },
        "name": "efs-input"
      },
      {
        "efsVolumeConfiguration": {
          "fileSystemId": "fs-你的EFS ID",
          "rootDirectory": "/output",
          "transitEncryption": "ENABLED"
        },
        "name": "efs-output"
      }
    ],
    "executionRoleArn": "arn:aws:iam::你的账号ID:role/BatchFargateExecutionRole",
    "jobRoleArn": "arn:aws:iam::你的账号ID:role/BatchJobS3AccessRole",
    "networkConfiguration": {
      "awsvpcConfiguration": {
        "subnets": ["subnet-你的子网ID"],
        "securityGroups": ["sg-你的安全组ID"],
        "assignPublicIp": "ENABLED"
      }
    }
  }
}
  • 需提前创建EFS文件系统,并确保Fargate任务所在子网能访问EFS。

关键注意事项

  • IAM权限:作业角色必须具备指定S3路径的读写权限;Fargate的执行角色需具备镜像拉取、日志写入等基础权限。
  • 大文件处理:若输入文件较大,优先选择EFS作为中间存储,避免重复下载。
  • 日志调试:可在作业定义中添加logConfiguration字段,将容器日志发送到CloudWatch,方便排查问题。

内容的提问来源于stack exchange,提问作者O.rka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 19:12:03