如何在AWS Batch作业定义JSON中指定输入输出S3卷路径?
AWS Batch作业定义配置指南(EC2/Fargate后端)
问题背景
我需要在AWS Batch上用EC2或Fargate后端运行VEBA预处理作业,本地Docker运行命令如下:
# Directories LOCAL_WORKING_DIRECTORY=$(pwd) LOCAL_OUTPUT_PARENT_DIRECTORY=../ LOCAL_OUTPUT_PARENT_DIRECTORY=$(realpath -m ${LOCAL_OUTPUT_PARENT_DIRECTORY}) CONTAINER_INPUT_DIRECTORY=/volumes/input/ CONTAINER_OUTPUT_DIRECTORY=/volumes/output/ # Parameters ID=S1 R1=Fastq/${ID}_1.fastq.gz R2=Fastq/${ID}_2.fastq.gz NAME=VEBA-preprocess__${ID} RELATIVE_OUTPUT_DIRECTORY=veba_output/preprocess/ # Command CMD="preprocess.py -1 ${CONTAINER_INPUT_DIRECTORY}/${R1} -2 ${CONTAINER_INPUT_DIRECTORY}/${R2} -n ${ID} -o ${CONTAINER_OUTPUT_DIRECTORY}/${RELATIVE_OUTPUT_DIRECTORY}" # Docker DOCKER_IMAGE="jolespin/veba_preprocess:1.1.2" docker run \ --name ${NAME} \ --rm \ --volume ${LOCAL_WORKING_DIRECTORY}:${CONTAINER_INPUT_DIRECTORY} \ --volume ${LOCAL_OUTPUT_PARENT_DIRECTORY}:${CONTAINER_OUTPUT_DIRECTORY} \ ${DOCKER_IMAGE} \ -c "${CMD}"
当前输入文件存于S3路径s3://path/to/input/(包含35_R1.fq.gz和35_R2.fq.gz),需将输出存储到s3://path/to/output/,但不清楚如何在AWS Batch作业定义JSON中配置S3路径的卷挂载,现有不完整的作业定义如下:
{ "jobDefinitionName": "preprocess__35", "type": "container", "containerProperties": { "image": "jolespin/veba_preprocess:1.1.2", "vcpus": 4, "memory": 16000, "command": [ "preprocess.py", "-1", "/volumes/input/35_R1.fq.gz", "-2", "/volumes/input/35_R2.fq.gz", "-n", "35", "-o", "/volumes/output/veba_output/preprocess", "-p", "4" ], "mountPoints": [ { } ], "volumes": [ ] } }
解决方案
AWS Batch本身不支持直接将S3作为卷挂载到容器,需根据后端类型(EC2/Fargate)选择合适的方式处理S3输入输出,以下是两种后端的完整配置:
一、EC2后端配置
EC2后端可选两种方案:利用实例本地存储+S3同步,或使用EFS持久化存储。
方案1:实例本地存储+S3同步(轻量快速)
直接在作业命令中完成S3输入下载、预处理、输出上传的流程,无需额外存储服务:
{ "jobDefinitionName": "preprocess__35", "type": "container", "platformCapabilities": ["EC2"], "containerProperties": { "image": "jolespin/veba_preprocess:1.1.2", "vcpus": 4, "memory": 16000, "command": [ "/bin/bash", "-c", "mkdir -p /volumes/input /volumes/output && aws s3 sync s3://path/to/input/ /volumes/input/ && preprocess.py -1 /volumes/input/35_R1.fq.gz -2 /volumes/input/35_R2.fq.gz -n 35 -o /volumes/output/veba_output/preprocess -p 4 && aws s3 sync /volumes/output/ s3://path/to/output/" ], "mountPoints": [], "volumes": [], "jobRoleArn": "arn:aws:iam::你的账号ID:role/BatchJobS3AccessRole" }, "retryStrategy": { "attempts": 1 } }
- 替换
jobRoleArn为你的IAM角色ARN,该角色需拥有s3:GetObject、s3:PutObject权限,允许访问指定S3路径。 - 用
/bin/bash -c串联命令:创建容器内目录、同步S3输入到容器、执行预处理、同步输出回S3。
方案2:挂载EC2实例本地卷
如果需要使用实例本地存储作为临时卷,可配置如下:
{ "jobDefinitionName": "preprocess__35", "type": "container", "platformCapabilities": ["EC2"], "containerProperties": { "image": "jolespin/veba_preprocess:1.1.2", "vcpus": 4, "memory": 16000, "command": [ "/bin/bash", "-c", "aws s3 sync s3://path/to/input/ /volumes/input/ && preprocess.py -1 /volumes/input/35_R1.fq.gz -2 /volumes/input/35_R2.fq.gz -n 35 -o /volumes/output/veba_output/preprocess -p 4 && aws s3 sync /volumes/output/ s3://path/to/output/" ], "mountPoints": [ { "sourceVolume": "input-volume", "containerPath": "/volumes/input", "readOnly": false }, { "sourceVolume": "output-volume", "containerPath": "/volumes/output", "readOnly": false } ], "volumes": [ { "host": { "sourcePath": "/tmp/input" }, "name": "input-volume" }, { "host": { "sourcePath": "/tmp/output" }, "name": "output-volume" } ], "jobRoleArn": "arn:aws:iam::你的账号ID:role/BatchJobS3AccessRole" } }
- 这里使用EC2实例的
/tmp目录作为卷,作业结束后会自动清理,适合临时存储。
二、Fargate后端配置
Fargate不支持实例本地卷,推荐使用S3同步方案,部分区域也支持挂载EFS。
方案1:S3同步(通用推荐)
配置方式类似EC2,但需添加Fargate专属的执行角色和网络配置:
{ "jobDefinitionName": "preprocess__35", "type": "container", "platformCapabilities": ["FARGATE"], "containerProperties": { "image": "jolespin/veba_preprocess:1.1.2", "vcpus": 4, "memory": 16000, "command": [ "/bin/bash", "-c", "mkdir -p /volumes/input /volumes/output && aws s3 sync s3://path/to/input/ /volumes/input/ && preprocess.py -1 /volumes/input/35_R1.fq.gz -2 /volumes/input/35_R2.fq.gz -n 35 -o /volumes/output/veba_output/preprocess -p 4 && aws s3 sync /volumes/output/ s3://path/to/output/" ], "mountPoints": [], "volumes": [], "executionRoleArn": "arn:aws:iam::你的账号ID:role/BatchFargateExecutionRole", "jobRoleArn": "arn:aws:iam::你的账号ID:role/BatchJobS3AccessRole", "networkConfiguration": { "awsvpcConfiguration": { "subnets": ["subnet-你的子网ID"], "securityGroups": ["sg-你的安全组ID"], "assignPublicIp": "ENABLED" } } }, "retryStrategy": { "attempts": 1 } }
executionRoleArn是Fargate执行角色,需拥有拉取镜像、写入CloudWatch日志的权限。networkConfiguration需配置你的VPC子网和安全组,assignPublicIp设为ENABLED让任务能访问S3(或使用VPC端点)。
方案2:挂载EFS(限支持区域)
若你的AWS区域支持Fargate挂载EFS,可先将S3内容同步到EFS,再挂载到容器:
{ "jobDefinitionName": "preprocess__35", "type": "container", "platformCapabilities": ["FARGATE"], "containerProperties": { "image": "jolespin/veba_preprocess:1.1.2", "vcpus": 4, "memory": 16000, "command": [ "/bin/bash", "-c", "aws s3 sync s3://path/to/input/ /volumes/input/ && preprocess.py -1 /volumes/input/35_R1.fq.gz -2 /volumes/input/35_R2.fq.gz -n 35 -o /volumes/output/veba_output/preprocess -p 4 && aws s3 sync /volumes/output/ s3://path/to/output/" ], "mountPoints": [ { "sourceVolume": "efs-input", "containerPath": "/volumes/input", "readOnly": false }, { "sourceVolume": "efs-output", "containerPath": "/volumes/output", "readOnly": false } ], "volumes": [ { "efsVolumeConfiguration": { "fileSystemId": "fs-你的EFS ID", "rootDirectory": "/input", "transitEncryption": "ENABLED" }, "name": "efs-input" }, { "efsVolumeConfiguration": { "fileSystemId": "fs-你的EFS ID", "rootDirectory": "/output", "transitEncryption": "ENABLED" }, "name": "efs-output" } ], "executionRoleArn": "arn:aws:iam::你的账号ID:role/BatchFargateExecutionRole", "jobRoleArn": "arn:aws:iam::你的账号ID:role/BatchJobS3AccessRole", "networkConfiguration": { "awsvpcConfiguration": { "subnets": ["subnet-你的子网ID"], "securityGroups": ["sg-你的安全组ID"], "assignPublicIp": "ENABLED" } } } }
- 需提前创建EFS文件系统,并确保Fargate任务所在子网能访问EFS。
关键注意事项
- IAM权限:作业角色必须具备指定S3路径的读写权限;Fargate的执行角色需具备镜像拉取、日志写入等基础权限。
- 大文件处理:若输入文件较大,优先选择EFS作为中间存储,避免重复下载。
- 日志调试:可在作业定义中添加
logConfiguration字段,将容器日志发送到CloudWatch,方便排查问题。
内容的提问来源于stack exchange,提问作者O.rka
相关产品推荐
相关产品推荐

