You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Prometheus的K8s服务级指标在Grafana中的展示需求

Got it, let's break down how to tackle both of your monitoring requirements using Prometheus and Grafana. I’ll walk you through the exact PromQL queries and setup tips for each scenario:

1. Monitor Service-Level Metrics for Pods Exposed via a Service

When you have 3 pods behind a Service, you’ll need to link the Service to its underlying pods using label selectors, then aggregate metrics across those pods. Here’s how you can query each key metric:

  • CPU Usage (Average per Pod, Aggregated for the Service)
    Use this query to get the average CPU usage across all pods tied to your target Service:

    avg(rate(container_cpu_usage_seconds_total{namespace="<your-namespace>", pod=~"$service-pod-pattern"}[5m])) by (pod)
    

    Note: Replace <your-namespace> with your K8s namespace, and $service-pod-pattern with a regex matching your Service's pods (e.g., my-app-.*). Alternatively, use the kube_service_spec_selector metric to dynamically link the Service to its pods:

    avg(rate(container_cpu_usage_seconds_total{namespace="<your-namespace>"}[5m])) by (pod)
    * on (namespace, label_app) group_left(service)
    kube_service_spec_selector{service="<your-service-name>", label_app="<your-app-label>"}
    
  • Memory Usage (Total & Per Pod)
    For total memory used by the Service's pods:

    sum(container_memory_working_set_bytes{namespace="<your-namespace>", pod=~"$service-pod-pattern"})
    

    For per-pod memory breakdown:

    container_memory_working_set_bytes{namespace="<your-namespace>", pod=~"$service-pod-pattern"}
    
  • Network IO (Bytes In/Out)
    Aggregated network incoming traffic for the Service:

    sum(rate(container_network_receive_bytes_total{namespace="<your-namespace>", pod=~"$service-pod-pattern"}[5m]))
    

    Aggregated network outgoing traffic:

    sum(rate(container_network_transmit_bytes_total{namespace="<your-namespace>", pod=~"$service-pod-pattern"}[5m]))
    
  • Total Requests & Failed Requests
    Assuming you’re using an ingress or service monitor that tracks request metrics (like NGINX Ingress or Prometheus Operator's ServiceMonitor), use these queries:
    Total requests:

    sum(rate(nginx_ingress_controller_requests{service="<your-service-name>"}[5m]))
    

    Failed requests (HTTP 4xx/5xx):

    sum(rate(nginx_ingress_controller_requests{service="<your-service-name>", status=~"4..|5.."}[5m]))
    

    If you’re using application-level metrics (like Spring Boot Actuator or Node.js Prometheus client), replace the metric name with your app’s request counter (e.g., http_requests_total).

2. Aggregate Metrics for Pods Without a Linked Service

For a group of pods belonging to the same application but not tied to a Service, you’ll use their shared application labels (like app=<your-app-name>) to aggregate metrics. Here’s how to build a single Grafana view for these pods:

  • Aggregated CPU Usage

    avg(rate(container_cpu_usage_seconds_total{namespace="<your-namespace>", label_app="<your-app-name>"}[5m])) by (pod)
    

    Use sum() instead of avg() if you want the total CPU consumed by all pods in the group.

  • Aggregated Memory Usage
    Total memory for the pod group:

    sum(container_memory_working_set_bytes{namespace="<your-namespace>", label_app="<your-app-name>"})
    

    Per-pod memory breakdown:

    container_memory_working_set_bytes{namespace="<your-namespace>", label_app="<your-app-name>"}
    
  • Aggregated Network IO
    Total incoming traffic:

    sum(rate(container_network_receive_bytes_total{namespace="<your-namespace>", label_app="<your-app-name>"}[5m]))
    

    Total outgoing traffic:

    sum(rate(container_network_transmit_bytes_total{namespace="<your-namespace>", label_app="<your-app-name>"}[5m]))
    
  • Application-Level Request Metrics
    If your pods expose request metrics, aggregate them using the app label:
    Total requests:

    sum(rate(http_requests_total{namespace="<your-namespace>", app="<your-app-name>"}[5m]))
    

    Failed requests:

    sum(rate(http_requests_total{namespace="<your-namespace>", app="<your-app-name>", status=~"4..|5.."}[5m]))
    

Grafana Setup Tips

  • Create variables in Grafana for namespace, service-name, and app-name so you can easily switch between different services or applications without rewriting queries.
  • For the unlinked pod group, add a panel that shows the list of pods (using kube_pod_info{label_app="<your-app-name>"}) to confirm which pods are included in the aggregation.
  • Use graph panels for time-series metrics (CPU, memory, network IO) and stat panels for total requests/failures to get a clear overview at a glance.

内容的提问来源于stack exchange,提问作者shiv455

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:52:13