Azure基础内部负载均衡器监控方案咨询
Great question—dealing with monitoring for Azure's Basic Internal Load Balancer (ILB) can feel limiting since it lacks the built-in telemetry of the Standard SKU, but there are several practical workarounds I’ve implemented in production environments that strike a balance between cost and visibility.
1. Leverage Azure Monitor with Custom Metrics & Logs
While the Basic ILB doesn’t expose native probe health metrics, you can proxy this data from your backend VMs using Azure Monitor:
- Deploy the Azure Monitor Agent on all backend VMs, then create a custom script extension to run periodic checks against your probe port (e.g.,
ss -tulnp | grep :<probe-port>for TCP probes). - Push the results as custom metrics to Azure Monitor using the
az monitor metricsCLI command or the Azure Monitor REST API. For example, a simple bash script snippet:# Check if probe port is listening PORT_LISTENING=$(ss -tulnp | grep :8080 | wc -l) # Push metric to Azure Monitor az monitor metrics create --resource-id <vm-resource-id> --namespace "Custom/ILBProbe" --name "ProbePortListening" --value $PORT_LISTENING --dimension "VMName=$(hostname)" - Enable VM Diagnostic Logs to collect systemd or syslog entries related to network connections (e.g., connection refusals on the probe port). Send these logs to Log Analytics, then create queries to flag anomalies (e.g., repeated
Connection refusedevents for your probe port).
2. Build a Lightweight Internal Monitoring Service
Create a dedicated monitoring node (or use an existing VM in the cluster) to simulate ILB health probe behavior:
- Prometheus + Grafana (Lightweight Deployment): Configure Prometheus to scrape custom metrics from each backend VM (e.g., using the
node_exporterwith custom scripts to expose probe port status). Set up alerts in Prometheus for failed probe checks, and use Grafana to build dashboards showing backend VM responsiveness over time. - Custom Scripted Checks: Write a Python or bash script that sends identical requests to your backend VMs as the ILB probe (e.g., HTTP GET to
/healthfor an HTTP probe). Schedule the script viacronand log results to a file or directly to Log Analytics. Example Python snippet for an HTTP probe:import requests from datetime import datetime backend_vms = ["vm-01.internal", "vm-02.internal"] probe_port = 8080 for vm in backend_vms: try: response = requests.get(f"http://{vm}:{probe_port}/health", timeout=5) status = "HEALTHY" if response.status_code == 200 else "UNHEALTHY" except requests.exceptions.RequestException as e: status = f"FAILED: {str(e)}" # Log to stdout (or send to Log Analytics via API) print(f"{datetime.utcnow()} | VM: {vm} | Status: {status}")
3. Use Azure Network Watcher for Targeted Diagnostics
Network Watcher offers low-cost tools to complement your monitoring:
- Connection Monitor: Create a test that mimics your ILB’s probe configuration (same port, protocol, and frequency) between a source VM (e.g., your monitoring node) and each backend VM. Network Watcher will log connectivity status, latency, and packet loss, and you can set up alerts for failures.
- Packet Capture: When troubleshooting specific issues, use Network Watcher to capture traffic to/from the ILB’s frontend IP or backend VMs. This avoids manual
tcpdumpsessions and lets you analyze captures directly in the Azure portal or download them for deeper inspection with tools like Wireshark.
4. Automate & Enhance Existing Tools
If you’re already using ping or tcpdump, make them more actionable:
- Automate Ping Checks: Schedule
pingtests viacronand pipe results to a log file. Use Log Analytics to ingest these logs and create alerts when packet loss exceeds a threshold (e.g., 20% over 5 minutes). - Filtered TCPDump: Instead of capturing all traffic, filter for your probe port to reduce noise. For example:
Pair this withtcpdump -i eth0 port 8080 -w probe_traffic.pcaptsharkto quickly analyze response codes:tshark -r probe_traffic.pcap -T fields -e ip.src -e tcp.flags -e http.response.code | grep -E "RST|500"
All these approaches avoid the cost of upgrading to the Standard ILB while giving you the visibility you need to diagnose probe failures and backend VM responsiveness issues.
内容的提问来源于stack exchange,提问作者shaswata pal

