You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何一次性读取100+订阅数据,构建VM CPU利用率监控仪表盘?

Great question—this is exactly the kind of automation that turns tedious daily email checks into proactive, actionable monitoring. Let’s break down how to build this scalable dashboard step by step, focusing on handling 100+ subscriptions and highlighting the most critical CPU utilization trends.

整体架构思路

We’ll structure this solution into three core, modular components to keep it maintainable and scalable:

  • Batch Data Collector: Pulls CPU metrics across all subscriptions and VMs
  • Data Processor: Cleans, aggregates, and filters for top CPU-consuming VMs
  • Visualization Dashboard: Renders trends in an interactive, easy-to-scan format
核心实现步骤

1. Batch Data Collection Across Subscriptions

First, you need a way to authenticate and fetch metrics across all your subscriptions without manual login. For cloud environments like Azure, using a service principal with monitoring read permissions is the way to go. Here’s a Python example using Azure’s SDKs:

from azure.identity import DefaultAzureCredential
from azure.mgmt.subscription import SubscriptionClient
from azure.mgmt.monitor import MonitorManagementClient
from azure.mgmt.compute import ComputeManagementClient
import pandas as pd

# Authenticate once for all subscriptions
credential = DefaultAzureCredential()
sub_client = SubscriptionClient(credential)

# Get all subscriptions
subscriptions = [sub for sub in sub_client.subscriptions.list()]

# Initialize empty dataframe to store metrics
all_metrics = pd.DataFrame()

for sub in subscriptions:
    sub_id = sub.subscription_id
    print(f"Fetching data for subscription: {sub.display_name}")
    
    # Initialize clients for this subscription
    monitor_client = MonitorManagementClient(credential, sub_id)
    compute_client = ComputeManagementClient(credential, sub_id)
    
    # Get all VMs in the subscription
    vms = [vm for vm in compute_client.virtual_machines.list_all()]
    
    # Fetch CPU metrics for each VM (last 24 hours, 1-hour intervals)
    for vm in vms:
        resource_id = vm.id
        metrics_data = monitor_client.metrics.list(
            resource_id,
            timespan="PT24H",
            interval="PT1H",
            metricnames="Percentage CPU",
            aggregation="Average"
        )
        
        # Parse metrics into dataframe
        for metric in metrics_data.value:
            for time_series in metric.timeseries:
                for data_point in time_series.data:
                    all_metrics = pd.concat([all_metrics, pd.DataFrame({
                        "Subscription": sub.display_name,
                        "VM Name": vm.name,
                        "Timestamp": data_point.time_stamp,
                        "CPU Utilization (%)": data_point.average
                    })])

2. Data Processing: Filter Top CPU VMs

Once you have all the raw data, you need to focus on the most critical VMs to avoid cluttering your dashboard. Use Pandas to aggregate and sort:

# Calculate average CPU per VM over the time period
vm_avg_cpu = all_metrics.groupby(["Subscription", "VM Name"])["CPU Utilization (%)"].mean().reset_index()

# Sort by CPU utilization and pick top N (e.g., top 20)
top_vms = vm_avg_cpu.sort_values(by="CPU Utilization (%)", ascending=False).head(20)

# Merge back with raw time-series data to get trends for top VMs
top_vm_trends = all_metrics.merge(top_vms[["Subscription", "VM Name"]], on=["Subscription", "VM Name"])

3. Build the Interactive Dashboard

For a quick, production-ready dashboard, use Streamlit—it’s perfect for building data apps with minimal code. Here’s a snippet to render the top VM CPU trends:

import streamlit as st
import plotly.express as px

# Set dashboard title
st.title("Top VM CPU Utilization Across Subscriptions")

# Add filters for subscription and time range
selected_sub = st.multiselect("Filter by Subscription", all_metrics["Subscription"].unique())
if selected_sub:
    filtered_data = top_vm_trends[top_vm_trends["Subscription"].isin(selected_sub)]
else:
    filtered_data = top_vm_trends

# Plot line chart for CPU trends
fig = px.line(filtered_data, 
              x="Timestamp", 
              y="CPU Utilization (%)", 
              color="VM Name",
              facet_col="Subscription",
              title="Hourly CPU Trend for Top VMs")

st.plotly_chart(fig, use_container_width=True)

# Display table of top VMs with average CPU
st.subheader("Top CPU-Consuming VMs (Average)")
st.dataframe(top_vms.style.format({"CPU Utilization (%)": "{:.2f}%"}))
Key Optimization Tips for 100+ Subscriptions
  • Parallelize Data Collection: Use asyncio or Python’s concurrent.futures to fetch data from multiple subscriptions at once—this cuts down runtime drastically.
  • Cache Reusable Data: Store VM resource IDs and subscription lists in a cache (like Redis) to avoid re-fetching them every time the dashboard runs.
  • Set Metric Retention: Only keep the last 7-30 days of data to keep your dataset manageable.
  • Add Threshold Alerts: Highlight VMs where CPU exceeds a predefined threshold (e.g., 90%) with red markers in the chart or a dedicated alert section.

内容的提问来源于stack exchange,提问作者Deepak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 17:57:53