You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Fargate集群部署多服务时容器报错‘Container is excited’无法启动咨询

Troubleshooting "Container is exited" Error for Fargate Services with 8080 Port Tasks

Got it, let's break down why your Fargate tasks are hitting that "Container is exited" error when deploying two services (each with 2 tasks running on port 8080) in the same cluster. First, quick note: I think that's a typo and you mean Container is exited—let's dive into the most likely fixes here:

1. Start with Checking Exit Codes & Logs (Non-Negotiable)

This is the first thing you should do every time a container exits unexpectedly. Head to the ECS Console:

  • Navigate to your cluster → the failing service → Tasks tab → select the stopped task → click the container name.
  • You’ll see the exit code (like 1, 137, 0) and can pull logs directly from CloudWatch here.
    • Exit code 1: Usually means your app crashed (missing config, invalid startup command, or uncaught errors).
    • Exit code 137: The container was killed because it exceeded allocated memory limits.
    • Exit code 0: The container ran its command and exited normally—if your app is supposed to be long-running, this means your startup command is wrong (e.g., it runs a one-off script instead of starting a server).

2. Port Conflicts? Probably Not (But Double-Check Mappings)

You might assume 8080 across tasks is the issue, but with Fargate’s default awsvpc network mode, every task gets its own elastic network interface (ENI) with a private IP. Port 8080 is isolated per task—no cross-task port conflicts here. That said, verify your task definition’s port settings:

  • Under the container configuration, make sure Container port is set to 8080.
  • If using a load balancer, leave Host port blank (Fargate handles this automatically). Setting it to 8080 isn’t harmful, but it’s unnecessary.

3. Resource Limits Are Too Low

If your app needs more CPU or memory than you’ve allocated in the task definition, the container will either fail to start or get killed shortly after.

  • For example, if you set 0.5 vCPU and 1GB memory but your app requires 2GB to run, it’ll crash with exit code 137. Try bumping up the resource allocation temporarily to test if that fixes the issue.

4. Invalid Startup Command or Broken Image

A common culprit is a misconfigured entrypoint/command in your task definition, or a Docker image that doesn’t work as expected.

  • First, test your image locally: Run docker run -p 8080:8080 your-image and confirm it starts up, listens on 8080, and stays running.
  • Then, cross-check your task definition’s Entry point and Command fields with what works locally. If your local run uses npm start, set the command to ["npm", "start"] (use JSON array format for Fargate).
  • Also, make sure your image includes all dependencies—don’t assume the base image has everything your app needs.

5. Health Check Failures Triggering Restarts

If you’ve set up health checks in your task definition, failing checks can cause ECS to kill and restart the container repeatedly, leading to a loop of exited tasks.

  • Verify your health check settings: For example, if your app’s health endpoint is /health on 8080, set the health check command to ["CMD-SHELL", "curl -f http://localhost:8080/health || exit 1"].
  • Adjust the interval and timeout if needed—if the health check runs too soon (before your app finishes starting), it’ll mark the container as unhealthy unnecessarily.

6. IAM Permissions or Network Issues

If your app needs to access AWS services (like ECR, S3) or external APIs, missing permissions or blocked network traffic can cause it to crash.

  • Ensure your task execution role has permissions to pull images from your container registry (e.g., ECR) and any other services your app uses.
  • Check your task’s security group: It should allow outbound traffic to the internet (if your app needs external access) or to other AWS resources (like RDS).

Quick Isolation Test

To narrow down the problem:

  • Deploy just one service first (e.g., Service1 with 2 tasks). If that works, the issue is specific to Service2’s task definition or image.
  • If even a single task fails, your problem is likely with the task definition or image itself—not the multi-service setup.

内容的提问来源于stack exchange,提问作者PriyaG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:58:23