如何在AWS ECS/Fargate上禁用Docker容器的核心转储文件
问题:AWS ECS/Fargate容器持续生成核心转储文件导致性能异常
我在AWS ECS/Fargate上运行基于ubuntu:22.04的Docker容器,搭载Python后端API服务(支持PDF等文档上传)。目前容器持续生成核心转储(core dump)文件,导致运行变慢甚至崩溃。
曾尝试在GitHub Actions的docker build命令中添加--ulimit core=0,但该参数仅适用于docker run场景;本地Docker的解决方案也无法适配ECS/Fargate生产环境,现寻求可行的禁用核心转储方案。
以下是当前的Dockerfile、ECS任务定义及GitHub Actions脚本:
现有Dockerfile配置
FROM python:3.9 RUN mkdir /code WORKDIR /code COPY requirements.txt . RUN pip install -r requirements.txt # Download the pandoc deb file RUN apt-get update && apt-get install -y wget RUN wget https://github.com/jgm/pandoc/releases/download/3.1.2/pandoc-3.1.2-1-amd64.deb # Install the downloaded deb file RUN dpkg -i pandoc-3.1.2-1-amd64.deb COPY . . CMD ["gunicorn", "-w", "17", "-k", "uvicorn.workers.UvicornWorker", "--timeout", "120", "main:app", "-b", "0.0.0.0:80"]
现有ECS任务定义
{ "taskDefinitionArn": "arn:aws:ecs:us-west-2:$ARN:task-definition/a$task-def:30", "containerDefinitions": [ { "name": "$NAME", "image": "$ARN.dkr.ecr.us-west-2.amazonaws.com/$IMAGE-NAME:22636912fe7ab73cf3bd23bdb3d88d317d00b272", "cpu": 0, "portMappings": [ { "name": "$CONTAINER_NAME-80-tcp", "containerPort": 80, "hostPort": 80, "protocol": "tcp", "appProtocol": "http" } ], "essential": true, "environment": [], "environmentFiles": [ { "value": "arn:aws:s3:::$S3_Resource", "type": "s3" } ], "mountPoints": [], "volumesFrom": [], "logConfiguration": { "logDriver": "awslogs", "options": { "awslogs-create-group": "true", "awslogs-group": "/ecs/$LOG_Group", "awslogs-region": "us-west-2", "awslogs-stream-prefix": "ecs" } } } ], "family": "$LOG_FAMILY", "taskRoleArn": "arn:aws:iam::$ARN:role/ecsTaskExecutionRole", "executionRoleArn": "arn:aws:iam::$ARN:role/ecsTaskExecutionRole", "networkMode": "awsvpc", "revision": 30, "volumes": [ { "name": "new", "host": {} } ], "status": "ACTIVE", "requiresAttributes": [ { "name": "com.amazonaws.ecs.capability.logging-driver.awslogs" }, { "name": "ecs.capability.execution-role-awslogs" }, { "name": "com.amazonaws.ecs.capability.ecr-auth" }, { "name": "com.amazonaws.ecs.capability.docker-remote-api.1.19" }, { "name": "ecs.capability.env-files.s3" }, { "name": "ecs.capability.increased-task-cpu-limit" }, { "name": "com.amazonaws.ecs.capability.task-iam-role" }, { "name": "ecs.capability.execution-role-ecr-pull" }, { "name": "ecs.capability.extensible-ephemeral-storage" }, { "name": "com.amazonaws.ecs.capability.docker-remote-api.1.18" }, { "name": "ecs.capability.task-eni" }, { "name": "com.amazonaws.ecs.capability.docker-remote-api.1.29" } ], "placementConstraints": [], "compatibilities": [ "EC2", "FARGATE" ], "requiresCompatibilities": [ "FARGATE" ], "cpu": "8192", "memory": "24576", "ephemeralStorage": { "sizeInGiB": 200 }, "runtimePlatform": { "cpuArchitecture": "X86_64", "operatingSystemFamily": "LINUX" }, "registeredAt": "2023-07-30T20:49:22.769Z", "registeredBy": "arn:aws:sts::$ARN:assumed-role/github/github", "tags": [] }
现有GitHub Actions脚本
name: Deploy Document-Management-service To Amazon ECS on: push: branches: - "main" env: AWS_REGION: # set this to preferred AWS region, e.g. us-west-1 ECR_REPOSITORY: # set this to your Amazon ECR repository name ECS_SERVICE: # set this to your Amazon ECS service name ECS_CLUSTER: # set this to your Amazon ECS cluster name ECS_TASK_DEFINITION: .github/workflows/main-task-definition.json # set this to the path to your Amazon ECS task definition # file, e.g. .aws/task-definition.json CONTAINER_NAME: # set this to the name of the container in the # containerDefinitions section of your task definition permissions: id-token: write contents: read # This is required for actions/checkout@v2 jobs: deploy: name: Deploy runs-on: ubuntu-latest environment: production steps: - name: Checkout uses: actions/checkout@v3 - name: Configure AWS credentials uses: aws-actions/configure-aws-credentials@v1 with: role-to-assume: ${{ secrets.AWS_ARN }} #AWS ARN With IAM Role role-session-name: github aws-region: ${{ env.AWS_REGION }} - name: Login to Amazon ECR id: login-ecr uses: aws-actions/amazon-ecr-login@v1 - name: Build, Push, Tag and Deploy Container to ECR. id: build-image env: ECR_REGISTRY: ${{ steps.login-ecr.outputs.registry }} IMAGE_TAG: ${{ github.sha }} run: | docker build -t $ECR_REGISTRY/$ECR_REPOSITORY:$IMAGE_TAG --ulimit core=0 . docker push $ECR_REGISTRY/$ECR_REPOSITORY:$IMAGE_TAG echo "::set-output name=image::$ECR_REGISTRY/$ECR_REPOSITORY:$IMAGE_TAG" - name: Fill in the new image ID in the Amazon ECS task definition id: task-def uses: aws-actions/amazon-ecs-render-task-definition@v1 with: task-definition: ${{ env.ECS_TASK_DEFINITION }} container-name: ${{ env.CONTAINER_NAME }} image: ${{ steps.build-image.outputs.image }} - name: Deploy Amazon ECS task definition uses: aws-actions/amazon-ecs-deploy-task-definition@v1 with: task-definition: ${{ steps.task-def.outputs.task-definition }} service: ${{ env.ECS_SERVICE }} cluster: ${{ env.ECS_CLUSTER }} wait-for-service-stability: true
可行解决方案
方案1:在Dockerfile中全局禁用核心转储
在镜像构建阶段配置系统级限制,确保容器启动后默认不生成core文件。
修改后的Dockerfile:
FROM python:3.9 # 全局禁用核心转储:设置软硬限制+清空core文件生成规则 RUN echo "* hard core 0" >> /etc/security/limits.conf && \ echo "* soft core 0" >> /etc/security/limits.conf && \ sysctl -w kernel.core_pattern= && \ echo "kernel.core_pattern=" >> /etc/sysctl.conf RUN mkdir /code WORKDIR /code COPY requirements.txt . RUN pip install -r requirements.txt # Download the pandoc deb file RUN apt-get update && apt-get install -y wget RUN wget https://github.com/jgm/pandoc/releases/download/3.1.2/pandoc-3.1.2-1-amd64.deb # Install the downloaded deb file RUN dpkg -i pandoc-3.1.2-1-amd64.deb COPY . . CMD ["gunicorn", "-w", "17", "-k", "uvicorn.workers.UvicornWorker", "--timeout", "120", "main:app", "-b", "0.0.0.0:80"]
方案2:在ECS任务定义中添加ulimit配置
直接在ECS容器定义中设置core文件的软硬限制,无需修改镜像。
修改任务定义的containerDefinitions节点,添加ulimits字段:
{ "containerDefinitions": [ { "name": "$NAME", "image": "$ARN.dkr.ecr.us-west-2.amazonaws.com/$IMAGE-NAME:22636912fe7ab73cf3bd23bdb3d88d317d00b272", "cpu": 0, "portMappings": [ { "name": "$CONTAINER_NAME-80-tcp", "containerPort": 80, "hostPort": 80, "protocol": "tcp", "appProtocol": "http" } ], "essential": true, "environment": [], "environmentFiles": [ { "value": "arn:aws:s3:::$S3_Resource", "type": "s3" } ], "mountPoints": [], "volumesFrom": [], "logConfiguration": { "logDriver": "awslogs", "options": { "awslogs-create-group": "true", "awslogs-group": "/ecs/$LOG_Group", "awslogs-region": "us-west-2", "awslogs-stream-prefix": "ecs" } }, // 添加ulimit配置 "ulimits": [ { "name": "core", "hardLimit": 0, "softLimit": 0 } ] } ], // 其他任务定义内容保持不变 }
方案3:在启动命令前临时禁用核心转储
修改Dockerfile的启动命令,在启动服务前执行ulimit -c 0临时禁用当前会话的core dump生成。
修改后的CMD部分:
CMD ["sh", "-c", "ulimit -c 0 && gunicorn -w 17 -k uvicorn.workers.UvicornWorker --timeout 120 main:app -b 0.0.0.0:80"]
验证方法
部署完成后,进入容器执行ulimit -c命令,输出0则表示禁用成功;同时监控容器内是否再生成core.*格式的文件。
内容的提问来源于stack exchange,提问作者h4nz0x
相关产品推荐
相关产品推荐

