如何通过AWS服务或脚本监控EC2实例端口并自动替换实例?
Great question! Let's break down your options here—CloudWatch can absolutely be part of the solution, and there are a few AWS-native and script-based approaches to handle this port monitoring + instance replacement workflow.
This is the most scalable, serverless approach. Here's how to set it up:
Deploy a port-check script on your EC2 instance
This script will regularly check if your target ports are listening, then send the status to CloudWatch as a custom metric. Usess(ornetstat) to verify listening ports, and the AWS CLI to push metrics.
Example script (save as/opt/port-monitor.shand make it executable):#!/bin/bash # Replace with your target ports TARGET_PORTS="80 443 8080" # Fetch instance metadata INSTANCE_ID=$(curl -s http://169.254.169.254/latest/meta-data/instance-id) REGION=$(curl -s http://169.254.169.254/latest/meta-data/placement/region) for PORT in $TARGET_PORTS; do # Check if port is listening (1 = listening, 0 = not) LISTENING_STATUS=$(ss -tuln | grep -E ":$PORT\s+" | wc -l) # Push metric to CloudWatch aws cloudwatch put-metric-data \ --region "$REGION" \ --namespace "EC2PortHealth" \ --metric-name "PortListening" \ --dimensions InstanceId="$INSTANCE_ID",Port="$PORT" \ --value "$LISTENING_STATUS" doneAdd it to cron to run every minute:
crontab -e # Add this line: * * * * * /opt/port-monitor.sh >> /var/log/port-monitor.log 2>&1Make sure your EC2 instance has an IAM role with the
cloudwatch:PutMetricDatapermission (restrict it to your metric namespace for security).Create CloudWatch Alarms for each port
For each port's custom metric, create an alarm that triggers when the value is0(not listening) for 1-2 minutes. Configure the alarm to invoke a Lambda function as its action.Build a Lambda function to replace the instance
This function will terminate the problematic instance and launch a new one from your Launch Template (critical for consistent instance configuration). Example Python code:import boto3 ec2_client = boto3.client('ec2') def lambda_handler(event, context): # Extract instance ID from the CloudWatch alarm event instance_id = event['detail']['dimensions']['InstanceId'] # Terminate the faulty instance ec2_client.terminate_instances(InstanceIds=[instance_id]) # Launch a new instance from your Launch Template ec2_client.run_instances( LaunchTemplate={ 'LaunchTemplateId': 'lt-xxxxxx', # Replace with your LT ID 'Version': '$Default' }, MinCount=1, MaxCount=1 ) return { 'statusCode': 200, 'body': f"Replaced instance {instance_id}: terminated old, launched new from Launch Template." }Assign an IAM role to the Lambda function with permissions for
ec2:TerminateInstancesandec2:RunInstances.
If you prefer a simpler setup without Lambda, you can run a monitoring script on a dedicated "watchdog" EC2 instance (or use EventBridge to trigger the script directly):
Write a port-check + instance-replacement script
Usencto test port connectivity, then use the AWS CLI to replace the instance if any port fails. Example:#!/bin/bash TARGET_INSTANCE_ID="i-xxxxxx" # Replace with your instance ID TARGET_PORTS="80 443 8080" REGION="us-east-1" LAUNCH_TEMPLATE_ID="lt-xxxxxx" # Get the target instance's private IP (adjust to public IP if needed) TARGET_IP=$(aws ec2 describe-instances --instance-ids "$TARGET_INSTANCE_ID" --region "$REGION" --query 'Reservations[0].Instances[0].PrivateIpAddress' --output text) # Check each port for PORT in $TARGET_PORTS; do nc -z -w 5 "$TARGET_IP" "$PORT" if [ $? -ne 0 ]; then echo "Port $PORT down on $TARGET_INSTANCE_ID — initiating replacement..." # Terminate old instance aws ec2 terminate-instances --instance-ids "$TARGET_INSTANCE_ID" --region "$REGION" # Launch new instance aws ec2 run-instances --launch-template LaunchTemplateId="$LAUNCH_TEMPLATE_ID" --region "$REGION" --min-count 1 --max-count 1 exit 0 fi done echo "All ports healthy on $TARGET_INSTANCE_ID."Schedule the script with EventBridge
Create an EventBridge rule to run the script every minute (using the "Run Command" action in Systems Manager, or by triggering the script on your watchdog instance).
If you want to avoid writing code entirely, use Systems Manager Automation to build a visual workflow:
- Create an Automation Document with steps:
- Use
aws:runCommandto execute a port-check script on the target EC2 instance. - Use
aws:executeAwsApito terminate the instance if ports are down. - Use
aws:executeAwsApito launch a new instance from your Launch Template.
- Use
- Trigger this Automation Document via a CloudWatch Alarm (when the custom port metric fails) or a scheduled EventBridge rule.
- Always use a Launch Template: This ensures your new instance has the same AMI, security groups, user data, and instance type as the old one.
- IAM Permissions: Double-check that all entities (EC2 instances, Lambda, Systems Manager) have the minimum required permissions to avoid security gaps.
- Add Notifications: Attach an SNS topic to your CloudWatch Alarms to get email/SMS alerts when an instance is replaced—critical for debugging.
- Auto Scaling Groups (Bonus): If you're running multiple instances, consider putting them in an Auto Scaling Group. You can configure CloudWatch Alarms to trigger ASG health checks and replace unhealthy instances automatically.
内容的提问来源于stack exchange,提问作者Cateno Viglio

