You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决低负载下AWS应用程序负载均衡器连接超时问题

Hey George, let's dig into why your AWS Application Load Balancer (ALB) isn't hitting that expected 600 req/sec and fix those connection timeouts. Here are the most common culprits and actionable fixes:

1. Backend Instance Connection Port Exhaustion

When testing directly to an instance, your client might reuse connections efficiently, but the ALB adds an extra layer that can lead to more short-lived connections (like TIME_WAIT states) piling up on your instances. If your instances run out of available ports to accept new connections from the ALB, you'll see timeouts even if the instance isn't at its 200 req/sec limit.

  • Check the issue: Run netstat -an | grep TIME_WAIT on your backend instances. If you see thousands of entries here, that's a red flag.
  • Fix it: Adjust your instance's kernel parameters to reuse TIME_WAIT connections:
    sysctl -w net.ipv4.tcp_tw_reuse=1
    sysctl -w net.ipv4.tcp_fin_timeout=30
    
    Persist these changes in /etc/sysctl.conf so they survive reboots. Also, confirm your ALB's target group has connection reuse enabled (it's on by default, but double-check in the AWS Console under Target Group > Attributes).
2. Misconfigured Target Group Health Checks

If your ALB's health checks are failing for one or more instances, those instances get taken out of the pool—meaning you're only using 1 or 2 instances instead of all 3. That immediately kills your scalability and can cause timeouts as remaining instances get overloaded.

  • Check the issue: Go to your Target Group in the AWS Console and look at the "Health status" column for all 3 instances. If any show unhealthy, that's the problem.
  • Fix it:
    • Verify your health check path returns a 200 OK response (test it directly via the instance's IP first).
    • Adjust health check timings: Set the Health check timeout to 2 seconds, Interval to 5 seconds, and Unhealthy threshold to 3. This gives instances enough time to respond without marking them down unnecessarily.
3. Mismatched Timeout Settings Between ALB and Backends

Your instance has a 5-second request timeout, but if your ALB's target timeout (the time it waits for a backend response) is shorter than that, the ALB will drop the connection before the instance can finish processing.

  • Check the issue: In your Target Group's Attributes, look for "Target timeout". If it's set to less than 5 seconds, that's the conflict.
  • Fix it: Increase the target timeout to 6 seconds (give a 1-second buffer over your instance's timeout) to ensure the ALB waits long enough for the instance to respond.
4. Test Tool Limitations

Sometimes the problem isn't with AWS—it's with your load testing tool. If your tool isn't configured to reuse connections (Keep-Alive) or hits its own resource limits (like client-side port exhaustion), it can't generate enough concurrent requests to reach 600 req/sec, leading to false timeouts.

  • Check the issue: If you're using ab (Apache Bench), for example, run it without -k and you'll see far more timeouts than with connection reuse.
  • Fix it: Use connection pooling in your test tool. For ab, add the -k flag:
    ab -n 10000 -c 600 -k http://your-alb-dns-name/your-test-path
    
    Also, ensure your test machine has enough CPU, memory, and available ports to handle the load (check with ulimit -n to see open file limits).
5. ALB DNS Load Balancing Not Working

ALBs use DNS to distribute traffic across multiple edge nodes. If your test client only resolves one ALB IP address, all traffic hits a single node, which might hit its own concurrency limits before your backend instances do.

  • Check the issue: Run nslookup your-alb-dns-name on your test machine. You should see multiple IP addresses returned. If only one shows up, your client is caching an old DNS entry.
  • Fix it: Flush your test machine's DNS cache, or test from a different client to ensure traffic is spread across ALB nodes.

If none of these fix the issue, enable ALB access logs (store them in an S3 bucket) and filter for 5xx errors or timeout responses. Look at the target_processing_time and target_status_code fields to see exactly which instances are causing timeouts and how long requests are taking.


内容的提问来源于stack exchange,提问作者George

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:00:34