基于跨节点部署的现有Prometheus实现AKS集群微服务监控及告警配置方案咨询
Hey Sharmila, let’s work through this problem step by step—you’re right, most guides assume Prometheus is in the same cluster, but we can make this work without touching your existing Prometheus setup. Here are two solid solutions tailored to your needs:
Solution 1: Deploy Prometheus Agent in AKS (No Changes to Existing Prometheus)
This is the cleanest approach if you want zero modifications to your current Prometheus configuration. We’ll use a lightweight Prometheus Agent in AKS to collect only the metrics we need, then send them directly to your existing Prometheus via remote write.
Step 1: Set Up Prometheus Agent in AKS
- Deploy the Prometheus Agent (a resource-light variant of Prometheus built for remote write) into your AKS cluster. You can use the official Helm chart for quick deployment.
- Configure the agent’s
remote_writesetting to point to your existing Prometheus’s remote write endpoint (usuallyhttp://<your-prometheus-ip>:9090/api/v1/write). - Configure the agent to scrape only the critical metrics for your use case:
- kube-state-metrics: This exposes pod restart counts (
kube_pod_container_status_restarts_total) and termination reasons (kube_pod_container_status_terminated_reason)—perfect for detecting crashes and OOM kills. - Kubelet metrics: Grab OOM-specific metrics like
kubelet_container_logs_oom_events_totaland memory usage metrics (container_memory_working_set_bytes) to spot impending OOM issues.
- kube-state-metrics: This exposes pod restart counts (
- Add relabel rules to filter out unnecessary metrics/labels—this keeps data transfer and storage overhead low.
Step 2: Verify Metric Ingestion
Once the agent is running, head to your existing Prometheus UI and query for the metrics mentioned above. If they appear, you’re good to go.
Step 3: Add Alert Rules
You can add these alert rules to your existing Prometheus (assuming you manage rules via separate files, which doesn’t require changing core Prometheus config):
Pod Unexpected Restart Alert
groups: - name: aks-pod-alerts rules: - alert: PodUnexpectedRestart expr: increase(kube_pod_container_status_restarts_total[15m]) > 2 for: 5m labels: severity: warning annotations: summary: Pod {{ $labels.pod }} ({{ $labels.namespace }}) restarted unexpectedly description: This pod has restarted {{ $value }} times in 15 minutes—check for application crashes or resource issues.
Pod OOM Kill Alert
groups: - name: aks-pod-alerts rules: - alert: PodOOMKilled expr: kube_pod_container_status_terminated_reason{reason="OOMKilled"} == 1 for: 1m labels: severity: critical annotations: summary: Pod {{ $labels.pod }} ({{ $labels.namespace }}) killed by OOM description: The pod was terminated due to out-of-memory errors. Review memory limits and application memory usage. - alert: PodHighMemoryUsage expr: (container_memory_working_set_bytes / container_spec_memory_limit_bytes) * 100 > 90 for: 10m labels: severity: warning annotations: summary: Pod {{ $labels.pod }} ({{ $labels.namespace }}) using high memory description: The pod is using {{ $value }}% of its allocated memory—this may lead to OOM kills soon.
Solution 2: Prometheus Federation (Minor Config Change)
If you’re open to a tiny tweak to your existing Prometheus config (just adding one scrape job), federation is a great option. We’ll deploy a small Prometheus instance in AKS to collect cluster metrics, then have your existing Prometheus pull only the metrics it needs from this instance.
Step 1: Deploy a Small Prometheus in AKS
Use Helm to deploy a minimal Prometheus instance in AKS, configured to scrape kube-state-metrics and kubelet metrics.
Step 2: Add a Federation Scrape Job to Existing Prometheus
Add this job to your existing Prometheus’s scrape config (this is the only change needed):
scrape_configs: - job_name: 'aks-federation' scrape_interval: 15s static_configs: - targets: ['<aks-prometheus-ip>:9090'] metrics_path: '/federate' params: 'match[]': - '{job=~"kube-state-metrics|kubelet"}'
This tells your existing Prometheus to pull only the relevant AKS metrics from the cluster’s Prometheus instance.
Step 3: Set Up Alerts
Use the same alert rules from Solution 1 to detect restarts and OOM events.
Key Notes
- Network Connectivity: Ensure your AKS cluster can reach your existing Prometheus’s remote write endpoint (or vice versa for federation). If they’re in separate private networks, you may need to set up a VPN or private link.
- RBAC Permissions: When deploying kube-state-metrics in AKS, make sure it has the necessary RBAC permissions to access Kubernetes API data about pods and containers.
- Metric Filtering: Always filter metrics to only what you need—this reduces bandwidth usage and keeps your Prometheus storage efficient.
内容的提问来源于stack exchange,提问作者Sharmi

