单查询合并多个Prometheus指标失败问题求助
问题分析与解决方案
核心问题
- 合并查询无结果:
kube_pod_status_reason是Pod级别的指标,返回结果中没有container标签,你在on(pod, container, namespace)中指定了container作为匹配标签,导致左右两侧没有共同的container标签,无法完成匹配,因此无返回结果。 - 重复系列错误:一个Pod可能对应多个
reason(比如同一Pod出现Evicted和NodeAffinity两种状态),同时一个Pod可能包含多个容器(对应左侧多个带container标签的系列),属于多对多匹配场景,Prometheus默认禁止这种匹配,因此报错。
解决方案
方案1:保留Reason维度,统计各容器因特定原因的重启次数
通过on(pod, namespace)完成匹配(这两个标签是两边共有的),用group_left(container)保留左侧的container标签,同时允许右侧一个Pod对应多个Reason的情况,最后按容器、命名空间、Reason聚合:
sum by(container, namespace, reason) ( increase(kube_pod_container_status_restarts_total{namespace!~"excluded-namespace.*"}[10m]) * on(pod, namespace) group_left(container) kube_pod_status_reason{reason=~"^(Evicted|Shutdown|NodeAffinity|NodeLost|UnexpectedAdmissionError)$"} )
方案2:仅统计因指定原因导致的总重启次数
如果不需要区分具体Reason,先对右侧指标按Pod和Namespace聚合(确保每个Pod+Namespace只有一个系列),再和左侧重启次数相乘,最后按容器、命名空间聚合:
sum by(container, namespace) ( increase(kube_pod_container_status_restarts_total{namespace!~"excluded-namespace.*"}[10m]) * on(pod, namespace) group_left(container) count by(pod, namespace) (kube_pod_status_reason{reason=~"^(Evicted|Shutdown|NodeAffinity|NodeLost|UnexpectedAdmissionError)$"}) )
内容的提问来源于stack exchange,提问作者DisplayName
相关产品推荐
相关产品推荐

