You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过AWS服务或脚本监控EC2实例端口并自动替换实例?

Great question! Let's break down your options here—CloudWatch can absolutely be part of the solution, and there are a few AWS-native and script-based approaches to handle this port monitoring + instance replacement workflow.

Option 1: CloudWatch Custom Metrics + Lambda Automation (AWS-Native Standard)

This is the most scalable, serverless approach. Here's how to set it up:

  1. Deploy a port-check script on your EC2 instance
    This script will regularly check if your target ports are listening, then send the status to CloudWatch as a custom metric. Use ss (or netstat) to verify listening ports, and the AWS CLI to push metrics.
    Example script (save as /opt/port-monitor.sh and make it executable):

    #!/bin/bash
    # Replace with your target ports
    TARGET_PORTS="80 443 8080"
    # Fetch instance metadata
    INSTANCE_ID=$(curl -s http://169.254.169.254/latest/meta-data/instance-id)
    REGION=$(curl -s http://169.254.169.254/latest/meta-data/placement/region)
    
    for PORT in $TARGET_PORTS; do
        # Check if port is listening (1 = listening, 0 = not)
        LISTENING_STATUS=$(ss -tuln | grep -E ":$PORT\s+" | wc -l)
        # Push metric to CloudWatch
        aws cloudwatch put-metric-data \
            --region "$REGION" \
            --namespace "EC2PortHealth" \
            --metric-name "PortListening" \
            --dimensions InstanceId="$INSTANCE_ID",Port="$PORT" \
            --value "$LISTENING_STATUS"
    done
    

    Add it to cron to run every minute:

    crontab -e
    # Add this line:
    * * * * * /opt/port-monitor.sh >> /var/log/port-monitor.log 2>&1
    

    Make sure your EC2 instance has an IAM role with the cloudwatch:PutMetricData permission (restrict it to your metric namespace for security).

  2. Create CloudWatch Alarms for each port
    For each port's custom metric, create an alarm that triggers when the value is 0 (not listening) for 1-2 minutes. Configure the alarm to invoke a Lambda function as its action.

  3. Build a Lambda function to replace the instance
    This function will terminate the problematic instance and launch a new one from your Launch Template (critical for consistent instance configuration). Example Python code:

    import boto3
    
    ec2_client = boto3.client('ec2')
    
    def lambda_handler(event, context):
        # Extract instance ID from the CloudWatch alarm event
        instance_id = event['detail']['dimensions']['InstanceId']
        
        # Terminate the faulty instance
        ec2_client.terminate_instances(InstanceIds=[instance_id])
        
        # Launch a new instance from your Launch Template
        ec2_client.run_instances(
            LaunchTemplate={
                'LaunchTemplateId': 'lt-xxxxxx',  # Replace with your LT ID
                'Version': '$Default'
            },
            MinCount=1,
            MaxCount=1
        )
        
        return {
            'statusCode': 200,
            'body': f"Replaced instance {instance_id}: terminated old, launched new from Launch Template."
        }
    

    Assign an IAM role to the Lambda function with permissions for ec2:TerminateInstances and ec2:RunInstances.

Option 2: Shell Script + EventBridge (Lightweight Script-Based)

If you prefer a simpler setup without Lambda, you can run a monitoring script on a dedicated "watchdog" EC2 instance (or use EventBridge to trigger the script directly):

  1. Write a port-check + instance-replacement script
    Use nc to test port connectivity, then use the AWS CLI to replace the instance if any port fails. Example:

    #!/bin/bash
    TARGET_INSTANCE_ID="i-xxxxxx"  # Replace with your instance ID
    TARGET_PORTS="80 443 8080"
    REGION="us-east-1"
    LAUNCH_TEMPLATE_ID="lt-xxxxxx"
    
    # Get the target instance's private IP (adjust to public IP if needed)
    TARGET_IP=$(aws ec2 describe-instances --instance-ids "$TARGET_INSTANCE_ID" --region "$REGION" --query 'Reservations[0].Instances[0].PrivateIpAddress' --output text)
    
    # Check each port
    for PORT in $TARGET_PORTS; do
        nc -z -w 5 "$TARGET_IP" "$PORT"
        if [ $? -ne 0 ]; then
            echo "Port $PORT down on $TARGET_INSTANCE_ID — initiating replacement..."
            # Terminate old instance
            aws ec2 terminate-instances --instance-ids "$TARGET_INSTANCE_ID" --region "$REGION"
            # Launch new instance
            aws ec2 run-instances --launch-template LaunchTemplateId="$LAUNCH_TEMPLATE_ID" --region "$REGION" --min-count 1 --max-count 1
            exit 0
        fi
    done
    echo "All ports healthy on $TARGET_INSTANCE_ID."
    
  2. Schedule the script with EventBridge
    Create an EventBridge rule to run the script every minute (using the "Run Command" action in Systems Manager, or by triggering the script on your watchdog instance).

Option 3: AWS Systems Manager Automation (No-Code Workflow)

If you want to avoid writing code entirely, use Systems Manager Automation to build a visual workflow:

  1. Create an Automation Document with steps:
    • Use aws:runCommand to execute a port-check script on the target EC2 instance.
    • Use aws:executeAwsApi to terminate the instance if ports are down.
    • Use aws:executeAwsApi to launch a new instance from your Launch Template.
  2. Trigger this Automation Document via a CloudWatch Alarm (when the custom port metric fails) or a scheduled EventBridge rule.
Key Notes to Remember
  • Always use a Launch Template: This ensures your new instance has the same AMI, security groups, user data, and instance type as the old one.
  • IAM Permissions: Double-check that all entities (EC2 instances, Lambda, Systems Manager) have the minimum required permissions to avoid security gaps.
  • Add Notifications: Attach an SNS topic to your CloudWatch Alarms to get email/SMS alerts when an instance is replaced—critical for debugging.
  • Auto Scaling Groups (Bonus): If you're running multiple instances, consider putting them in an Auto Scaling Group. You can configure CloudWatch Alarms to trigger ASG health checks and replace unhealthy instances automatically.

内容的提问来源于stack exchange,提问作者Cateno Viglio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:29:04