非Web应用服务器的AWS ALB简易健康检查方案咨询
Alright, let's tackle your two questions step by step—first configuring AWS ALB health checks for your backend servers, then a simple solution for those non-web servers without HTTP services.
If your backend servers run web services (HTTP/HTTPS), here's how to set up reliable health checks via the AWS Console or CLI:
Console Setup
- Navigate to the EC2 Dashboard → Target Groups and select the target group linked to your ALB.
- Click Edit health check to adjust these key settings:
- Protocol: Pick HTTP or HTTPS to match your server's service.
- Path: Use a dedicated health endpoint like
/health(preferred) or a simple path like/—ensure your server returns a 200-399 status code for this path. - Port: Choose "Traffic port" to use the same port as your application, or specify a fixed port if needed.
- Thresholds: Set
Healthy threshold count(e.g., 2 successful checks to mark as healthy) andUnhealthy threshold count(e.g., 3 failed checks to mark as unhealthy). - Timings: Adjust
Health check timeout(e.g., 5 seconds) andHealth check interval(e.g., 10 seconds) based on your server's response speed. - Success codes: Define valid HTTP status codes (default is 200, but you can expand to 200-399 if needed).
- Save the changes, and the ALB will start periodically verifying your servers' health—unhealthy instances get removed from traffic rotation automatically.
CLI Setup
If you prefer command-line configuration, use this example (replace placeholders with your actual values):
aws elbv2 modify-target-group \ --target-group-arn arn:aws:elasticloadbalancing:us-east-1:123456789012:targetgroup/my-target-group/abc123 \ --health-check-protocol HTTP \ --health-check-port traffic-port \ --health-check-path /health \ --healthy-threshold-count 2 \ --unhealthy-threshold-count 3 \ --health-check-timeout-seconds 5 \ --health-check-interval-seconds 10 \ --matcher HttpCode=200
Since your servers don't have HTTP services, we'll focus on lightweight, no-fuss options:
Option 1: TCP Port Health Check (Easiest)
ALB supports TCP-based health checks, which only require your server to listen on a specific port—no extra software needed.
- On your non-web server, set up a persistent TCP listener. For example, use
nc(netcat) with a systemd service to keep it running:- Create a service file
/etc/systemd/system/tcp-healthcheck.service:[Unit] Description=TCP Health Check Listener After=network.target [Service] ExecStart=/usr/bin/nc -l 10000 Restart=always [Install] WantedBy=multi-user.target - Enable and start the service:
sudo systemctl daemon-reload sudo systemctl enable --now tcp-healthcheck.service
- Create a service file
- Back in your ALB target group, edit the health check to use:
- Protocol: TCP
- Port: 10000 (or whatever port you chose)
- Adjust thresholds and timings as needed.
- The ALB will attempt to establish a TCP connection to this port—successful connections mean the server is healthy; failed connections trigger removal from the target group.
Option 2: Custom Health Checks (For Deep Verification)
If you need to check more than just an open port (e.g., if a specific process is running, disk space is sufficient), use a script + CloudWatch:
- Create a simple health check script (
/usr/local/bin/health-check.sh):#!/bin/bash # Check if your non-web service process is running if pgrep "my-critical-service" > /dev/null; then # Report healthy status to CloudWatch aws cloudwatch put-metric-data --namespace "NonWebServers" --metric-name "HealthStatus" --value 1 --dimensions InstanceId=$(curl -s http://169.254.169.254/latest/meta-data/instance-id) exit 0 else # Report unhealthy status aws cloudwatch put-metric-data --namespace "NonWebServers" --metric-name "HealthStatus" --value 0 --dimensions InstanceId=$(curl -s http://169.254.169.254/latest/meta-data/instance-id) exit 1 fi - Make it executable:
sudo chmod +x /usr/local/bin/health-check.sh - Set up a cron job to run it every minute:
* * * * * /usr/local/bin/health-check.sh - Create a CloudWatch Alarm that triggers when
HealthStatusis 0 for 2 consecutive minutes. Link this alarm to a Lambda function that removes the unhealthy instance from your ALB target group.
This approach gives you more control over what defines a "healthy" server, beyond just an open port.
内容的提问来源于stack exchange,提问作者Rohit Sharma

