请求实现磁盘占比超阈值时从Load Balancer移除实例的方案
Nice question—this is a super common reliability pattern to prevent degraded instances from serving traffic. I’ve implemented similar solutions for several production environments, so let’s walk through a solid, script-based approach.
The workflow is straightforward:
- Monitor disk usage at regular intervals
- Trigger an action if usage exceeds your defined threshold (e.g., 90%)
- Verify if the instance is still registered with your Load Balancer (LB)
- Remove the instance from the LB if it’s active, plus send alerts for visibility
First, create a script that checks your target disk’s usage. This works on any Linux-based instance:
#!/bin/bash # -------------------------- Configuration -------------------------- # Set your disk usage threshold (e.g., 90% = 90) DISK_THRESHOLD=90 # Target mount point (adjust to /data if using a separate data disk) TARGET_DISK="/" # ------------------------------------------------------------------- # Get current disk usage percentage (strip the % symbol) CURRENT_USAGE=$(df -h "$TARGET_DISK" | grep -v Filesystem | awk '{print $5}' | sed 's/%//g') echo "[$(date)] Current disk usage for $TARGET_DISK: $CURRENT_USAGE%" >> /var/log/disk-lb-monitor.log if [ "$CURRENT_USAGE" -ge "$DISK_THRESHOLD" ]; then echo "[$(date)] Disk usage exceeds threshold! Initiating LB removal..." >> /var/log/disk-lb-monitor.log # Call LB removal function (defined below) remove_from_lb else echo "[$(date)] Disk usage is within acceptable limits." >> /var/log/disk-lb-monitor.log fi
Key Notes:
- Adjust
DISK_THRESHOLDto match your team’s tolerance (common values are 85-95%) - Update
TARGET_DISKif you’re monitoring a non-root disk (e.g.,/mnt/data) - Logs are written to
/var/log/disk-lb-monitor.logfor debugging
Below are examples for major cloud providers. Pick the one that matches your environment, and add it as a function to the script above.
Example 1: AWS (EC2 + Classic/Application Load Balancer)
First, ensure your instance has an IAM role with permissions:
elasticloadbalancing:DeregisterInstancesFromLoadBalancer(for Classic LB)elasticloadbalancing:DeregisterTargets(for Application/Network LB)
Add this function to your script:
remove_from_lb() { # Get current EC2 instance ID via metadata service INSTANCE_ID=$(curl -s http://169.254.169.254/latest/meta-data/instance-id) # Replace with your LB name (Classic) or Target Group ARN (ALB/NLB) LB_RESOURCE="your-target-group-arn" # Check if instance is still registered REGISTERED=$(aws elbv2 describe-target-health --target-group-arn "$LB_RESOURCE" --targets Id=$INSTANCE_ID | grep -c "healthy") if [ "$REGISTERED" -gt 0 ]; then # Deregister from ALB/NLB (use aws elb deregister-instances-from-load-balancer for Classic LB) aws elbv2 deregister-targets --target-group-arn "$LB_RESOURCE" --targets Id=$INSTANCE_ID echo "[$(date)] Successfully deregistered instance $INSTANCE_ID from LB resource $LB_RESOURCE" >> /var/log/disk-lb-monitor.log # Optional: Send alert (email/Slack) echo "Disk Full Alert: Instance $INSTANCE_ID removed from LB. Usage: $CURRENT_USAGE%" | mail -s "Critical: Disk Full on EC2 Instance" ops@yourcompany.com else echo "[$(date)] Instance $INSTANCE_ID is not registered with LB resource $LB_RESOURCE" >> /var/log/disk-lb-monitor.log fi }
Example 2: GCP (Compute Engine + Backend Service)
Ensure your instance has the compute.backendServices.update permission. Add this function:
remove_from_lb() { # Get current GCE instance ID INSTANCE_ID=$(curl -s http://metadata.google.internal/computeMetadata/v1/instance/id -H "Metadata-Flavor: Google") # Replace with your backend service name and region BACKEND_SERVICE="your-backend-service" REGION="us-central1" # Deregister instance from backend service gcloud compute backend-services remove-backend "$BACKEND_SERVICE" \ --region="$REGION" \ --instance="$INSTANCE_ID" \ --zone=$(curl -s http://metadata.google.internal/computeMetadata/v1/instance/zone -H "Metadata-Flavor: Google" | cut -d/ -f4) echo "[$(date)] Successfully removed instance $INSTANCE_ID from backend service $BACKEND_SERVICE" >> /var/log/disk-lb-monitor.log }
To run the script every 5 minutes (adjust interval as needed), add a cron job:
- Open crontab for editing:
crontab -e - Add this line (update the script path to match where you saved it):
*/5 * * * * /usr/local/scripts/disk-lb-monitor.sh
- Add Re-Registration Logic: Extend the script to re-add the instance to the LB once disk usage drops below the threshold (e.g., 70%)
- Centralized Alerts: Replace email alerts with tools like PagerDuty, Slack, or Opsgenie for faster response
- Test the Workflow: Simulate a full disk with
dd if=/dev/zero of=/tmp/bigfile bs=1G count=10to verify the script triggers correctly - Limit Permissions: Use the principle of least privilege for IAM roles—don’t grant full admin access
- Monitor Logs: Set up log aggregation (e.g., CloudWatch, Stackdriver) to track script runs and LB changes
内容的提问来源于stack exchange,提问作者hamdi Hasnaoui

