You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求实现磁盘占比超阈值时从Load Balancer移除实例的方案

Nice question—this is a super common reliability pattern to prevent degraded instances from serving traffic. I’ve implemented similar solutions for several production environments, so let’s walk through a solid, script-based approach.

Core Solution Overview

The workflow is straightforward:

  • Monitor disk usage at regular intervals
  • Trigger an action if usage exceeds your defined threshold (e.g., 90%)
  • Verify if the instance is still registered with your Load Balancer (LB)
  • Remove the instance from the LB if it’s active, plus send alerts for visibility

Step 1: Disk Usage Monitoring Script (Base Logic)

First, create a script that checks your target disk’s usage. This works on any Linux-based instance:

#!/bin/bash

# -------------------------- Configuration --------------------------
# Set your disk usage threshold (e.g., 90% = 90)
DISK_THRESHOLD=90
# Target mount point (adjust to /data if using a separate data disk)
TARGET_DISK="/"
# -------------------------------------------------------------------

# Get current disk usage percentage (strip the % symbol)
CURRENT_USAGE=$(df -h "$TARGET_DISK" | grep -v Filesystem | awk '{print $5}' | sed 's/%//g')

echo "[$(date)] Current disk usage for $TARGET_DISK: $CURRENT_USAGE%" >> /var/log/disk-lb-monitor.log

if [ "$CURRENT_USAGE" -ge "$DISK_THRESHOLD" ]; then
    echo "[$(date)] Disk usage exceeds threshold! Initiating LB removal..." >> /var/log/disk-lb-monitor.log
    # Call LB removal function (defined below)
    remove_from_lb
else
    echo "[$(date)] Disk usage is within acceptable limits." >> /var/log/disk-lb-monitor.log
fi

Key Notes:

  • Adjust DISK_THRESHOLD to match your team’s tolerance (common values are 85-95%)
  • Update TARGET_DISK if you’re monitoring a non-root disk (e.g., /mnt/data)
  • Logs are written to /var/log/disk-lb-monitor.log for debugging

Step 2: Add Load Balancer Removal Logic (Cloud-Specific)

Below are examples for major cloud providers. Pick the one that matches your environment, and add it as a function to the script above.

Example 1: AWS (EC2 + Classic/Application Load Balancer)

First, ensure your instance has an IAM role with permissions:

  • elasticloadbalancing:DeregisterInstancesFromLoadBalancer (for Classic LB)
  • elasticloadbalancing:DeregisterTargets (for Application/Network LB)

Add this function to your script:

remove_from_lb() {
    # Get current EC2 instance ID via metadata service
    INSTANCE_ID=$(curl -s http://169.254.169.254/latest/meta-data/instance-id)
    # Replace with your LB name (Classic) or Target Group ARN (ALB/NLB)
    LB_RESOURCE="your-target-group-arn"

    # Check if instance is still registered
    REGISTERED=$(aws elbv2 describe-target-health --target-group-arn "$LB_RESOURCE" --targets Id=$INSTANCE_ID | grep -c "healthy")

    if [ "$REGISTERED" -gt 0 ]; then
        # Deregister from ALB/NLB (use aws elb deregister-instances-from-load-balancer for Classic LB)
        aws elbv2 deregister-targets --target-group-arn "$LB_RESOURCE" --targets Id=$INSTANCE_ID
        echo "[$(date)] Successfully deregistered instance $INSTANCE_ID from LB resource $LB_RESOURCE" >> /var/log/disk-lb-monitor.log
        # Optional: Send alert (email/Slack)
        echo "Disk Full Alert: Instance $INSTANCE_ID removed from LB. Usage: $CURRENT_USAGE%" | mail -s "Critical: Disk Full on EC2 Instance" ops@yourcompany.com
    else
        echo "[$(date)] Instance $INSTANCE_ID is not registered with LB resource $LB_RESOURCE" >> /var/log/disk-lb-monitor.log
    fi
}

Example 2: GCP (Compute Engine + Backend Service)

Ensure your instance has the compute.backendServices.update permission. Add this function:

remove_from_lb() {
    # Get current GCE instance ID
    INSTANCE_ID=$(curl -s http://metadata.google.internal/computeMetadata/v1/instance/id -H "Metadata-Flavor: Google")
    # Replace with your backend service name and region
    BACKEND_SERVICE="your-backend-service"
    REGION="us-central1"

    # Deregister instance from backend service
    gcloud compute backend-services remove-backend "$BACKEND_SERVICE" \
        --region="$REGION" \
        --instance="$INSTANCE_ID" \
        --zone=$(curl -s http://metadata.google.internal/computeMetadata/v1/instance/zone -H "Metadata-Flavor: Google" | cut -d/ -f4)
    
    echo "[$(date)] Successfully removed instance $INSTANCE_ID from backend service $BACKEND_SERVICE" >> /var/log/disk-lb-monitor.log
}

Step 3: Automate with Cron

To run the script every 5 minutes (adjust interval as needed), add a cron job:

  1. Open crontab for editing:
    crontab -e
    
  2. Add this line (update the script path to match where you saved it):
    */5 * * * * /usr/local/scripts/disk-lb-monitor.sh
    

Best Practices for Production
  • Add Re-Registration Logic: Extend the script to re-add the instance to the LB once disk usage drops below the threshold (e.g., 70%)
  • Centralized Alerts: Replace email alerts with tools like PagerDuty, Slack, or Opsgenie for faster response
  • Test the Workflow: Simulate a full disk with dd if=/dev/zero of=/tmp/bigfile bs=1G count=10 to verify the script triggers correctly
  • Limit Permissions: Use the principle of least privilege for IAM roles—don’t grant full admin access
  • Monitor Logs: Set up log aggregation (e.g., CloudWatch, Stackdriver) to track script runs and LB changes

内容的提问来源于stack exchange,提问作者hamdi Hasnaoui

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:57:34