AWS ECS结合ALB构建高可用应用遇流量瓶颈求解
Let's break down your problem first: you're hitting a throughput bottleneck at ~3k requests/min with minimal resource utilization on both ECS instances and ALB, and timeouts kick in before auto-scaling even triggers. This is almost certainly not an instance performance or ALB capacity issue—instead, we need to dig into configuration gaps, traffic path constraints, or load testing setup.
First: Do You Need an NLB in Front of ALB?
Short answer: No, not for your current scenario. NLBs are useful for extreme throughput (100k+ requests/sec), cross-region traffic routing, or TCP/UDP protocol support. Your ALB is barely hitting 0.27 CU, which is a tiny fraction of its capacity. Adding an NLB now would just add unnecessary complexity without addressing the root cause.
Step-by-Step Troubleshooting & Fixes
1. Verify Auto-Scaling Configuration is Actually Triggerable
Your service isn't scaling, which means either your alarms aren't firing, or the scaling policy isn't properly linked:
- Check CloudWatch Alarm thresholds: If you set CPU scaling to trigger at 70% utilization, but your containers only hit 5%, it will never fire. Lower the threshold temporarily (e.g., 20% CPU) and re-run your load test to see if scaling kicks in.
- Confirm metric sources: Ensure your alarm is using the
ECS Service CPUUtilizationmetric (not the underlying EC2 instance CPU). Instance CPU might be low because your tasks are underutilized, but the service-level metric is what drives scaling. - Check service limits: Verify your ECS service's
Maximum Task Countand the linked Auto Scaling Group'sMax Sizearen't set to a low number (e.g., 2 tasks max). If the service can't scale out, throughput will cap.
2. Optimize Load Testing Setup
loader.io's configuration might be limiting how much traffic it can send to your ALB:
- Increase concurrent connections: If you're using a low number of concurrent threads/users, you'll hit a ceiling on requests per minute even if your infrastructure can handle more. Try doubling or tripling the concurrency.
- Enable keep-alive: For HTTP/1.1 requests, ensure keep-alive is enabled to reuse TCP connections. Re-establishing connections for every request adds massive overhead and limits throughput.
- Use a test region close to your AWS resources: If loader.io's test nodes are in a different region, network latency will cause timeouts before your infrastructure can process requests. Pick a test region matching your AWS deployment.
3. Audit ALB & Target Group Configuration
Misconfigured ALB settings can block traffic before it reaches your tasks:
- Check ALB idle timeout: The default is 60 seconds, but if your load test uses short-lived connections, ensure this is set appropriately. If connections are dropped too early, you'll see timeouts.
- Review target group health checks: Overly aggressive health checks (e.g., 5-second intervals with 2-second timeouts) can cause the ALB to mark tasks as unhealthy and stop routing traffic to them. Ensure health check thresholds align with your task's response time.
- Analyze ALB access logs: Enable ALB access logs and look at the
elb_status_codeandtarget_status_codefields. If you see504 Gateway Timeout, that means the ALB waited too long for a response from your task. If you see client-side timeouts, the issue is likely with the load test tool or network path to the ALB.
4. Check VPC & Network Constraints
Hidden network limits can throttle traffic without showing up in resource metrics:
- Security Groups & Network ACLs: Ensure your ALB's security group allows enough inbound connections (no rate limits on the listener port) and that your ECS task security groups allow inbound traffic from the ALB.
- ENI limits: Each EC2 instance has a limit on elastic network interfaces and concurrent connections. For t2/t3 instances, this is usually in the thousands, so it's unlikely to be the issue here—but it's worth confirming if you're running a large number of tasks per instance.
5. Validate Task Definition Resources
Even if CPU utilization is low, your task's resource limits might be restricting throughput:
- Check CPU/memory reservations: If you allocated only 0.1 vCPU per task, even at 5% utilization, that's just 0.005 vCPU available per task—enough for a simple endpoint, but maybe not enough to handle concurrent requests efficiently. Try increasing the CPU reservation slightly (e.g., 0.2 vCPU) to see if throughput improves.
Final Recommendation
Start with the auto-scaling and load test configuration checks first—those are the most likely culprits. Once you get auto-scaling working and optimize your load test, you should see throughput increase significantly. Only consider adding an NLB if you later hit the ALB's hard throughput limits (which is way higher than your current 3k requests/min).
内容的提问来源于stack exchange,提问作者rorypicko

