异构基础设施监控方案推荐及Prometheus集成疑问咨询
Great choice picking Prometheus for your monitoring setup—it aligns perfectly with all your requirements, and integrating your custom Python script for external service metrics is straightforward. Let’s walk through how to make this work:
You’re on the right track here:
- Use wmi_exporter for Windows VMs to pull CPU, memory, disk I/O, and free space metrics.
- For Linux VMs, node_exporter covers the same base resource metrics.
- sql_exporter is ideal for SQL Server—you can define custom queries in its config to extract exactly the metrics you need (like query latency, connection counts, or database size). Map those queries to Prometheus-compatible metrics, and it’ll expose them for scraping automatically.
This is the core question you have, and there are two reliable ways to hook your script into Prometheus:
Option 1: Expose Metrics via an HTTP Endpoint (For Long-Running Scripts)
If your Python script runs continuously (polling the external service at regular intervals), use the prometheus-client library to expose metrics on an HTTP endpoint that Prometheus can scrape directly.
Here’s a simplified example:
from prometheus_client import start_http_server, Gauge import time import requests # Use this to call your external service's API # Define Prometheus gauges to track your metrics ACTIVE_TASKS = Gauge('external_service_active_tasks', 'Number of currently running tasks in external service') AVG_TASK_DURATION = Gauge('external_service_avg_task_duration_seconds', 'Average duration of recent tasks in external service') def fetch_and_update_metrics(): # Replace this with your logic to pull data from the external service service_response = requests.get('https://your-external-service/api/task-status') task_data = service_response.json() # Update the gauges with fresh data ACTIVE_TASKS.set(task_data['active_task_count']) AVG_TASK_DURATION.set(task_data['average_task_duration']) if __name__ == '__main__': # Start an HTTP server on port 8000 to expose metrics start_http_server(8000) # Poll the external service every 60 seconds while True: fetch_and_update_metrics() time.sleep(60)
Once the script is running, add a scrape job to your Prometheus config:
scrape_configs: - job_name: 'external_service_python' static_configs: - targets: ['your-script-host:8000']
Option 2: Push Metrics to Pushgateway (For Scheduled Scripts)
If your script runs on a schedule (e.g., via cron) instead of continuously, use Pushgateway to temporarily store metrics so Prometheus can scrape them.
First, deploy Pushgateway and add it to your Prometheus scrape config. Then modify your script to push metrics:
from prometheus_client import CollectorRegistry, Gauge, push_to_gateway import requests registry = CollectorRegistry() ACTIVE_TASKS = Gauge('external_service_active_tasks', 'Number of currently running tasks', registry=registry) AVG_TASK_DURATION = Gauge('external_service_avg_task_duration_seconds', 'Average task duration', registry=registry) # Fetch data from your external service service_response = requests.get('https://your-external-service/api/task-status') task_data = service_response.json() # Set metric values ACTIVE_TASKS.set(task_data['active_task_count']) AVG_TASK_DURATION.set(task_data['average_task_duration']) # Push metrics to Pushgateway (replace with your Pushgateway address) push_to_gateway('pushgateway:9091', job='external_service', registry=registry)
Run this script on your desired schedule, and Prometheus will pull the metrics from Pushgateway automatically.
- Charting: Use Grafana to build custom dashboards. Connect it to your Prometheus data source, then add panels for each metric (VM resource usage, SQL Server stats, external service task metrics). Grafana has pre-built Prometheus templates you can tweak to fit your needs.
- Alerting: Set up alert rules in Prometheus to trigger when thresholds are hit (e.g., CPU usage > 90%, task duration > 30 minutes). Then configure Alertmanager to send notifications:
- For email: Use the
email_configsreceiver in Alertmanager’s config. - For Slack/MS Teams: Use their respective webhook integrations—Alertmanager supports both natively with simple config snippets.
- For email: Use the
- Deploy wmi_exporter/node_exporter on your VMs, and sql_exporter for SQL Server.
- Run your Python script (either as a long-running service exposing an endpoint, or scheduled to push metrics to Pushgateway).
- Configure Prometheus to scrape all these sources.
- Build Grafana dashboards for real-time visualization.
- Set up Prometheus alert rules and Alertmanager for multi-channel notifications.
内容的提问来源于stack exchange,提问作者HunOL

