Podmetrics无法被prometheus-operator正常采集抓取问题排查
问题描述
我当前正在为示例应用配置PodMetrics功能。
应用已成功部署,以下是我当前使用的deployment.yaml配置文件:
apiVersion: apps/v1 kind: Deployment metadata: labels: app.kubernetes.io/name: prometheus-example-app name: prometheus-example-app spec: replicas: 1 selector: matchLabels: app.kubernetes.io/name: prometheus-example-app template: metadata: labels: app.kubernetes.io/name: prometheus-example-app prometheus.io/scrape: 'true' spec: containers: - name: prometheus-example-app image: quay.io/brancz/prometheus-example-app:v0.3.0 ports: - name: web containerPort: 8080
以下是我使用的Podmonitor.yaml配置文件:
apiVersion: monitoring.coreos.com/v1 kind: PodMonitor metadata: labels: app.kubernetes.io/name: prometheus-example-app name: prometheus-example-app spec: selector: matchLabels: app.kubernetes.io/name: prometheus-example-app podMetricsEndpoints: - port: web
我的prometheus-operator中未显示任何采集目标,以下是我部署prometheus operator使用的配置文件:
apiVersion: monitoring.coreos.com/v1 kind: Prometheus metadata: annotations: argocd.argoproj.io/sync-wave: "1" name: prometheus labels: name: prometheus spec: serviceAccountName: prometheus serviceMonitorSelector: {} serviceMonitorNamespaceSelector: matchLabels: prometheus-scrape: "true" podMonitorSelector: matchLabels: app.kubernetes.io/name: prometheus-example-app resources: requests: memory: 400Mi enableAdminAPI: false additionalScrapeConfigs: name: additional-scrape-configs key: prometheus-additional.yaml
我目前无法在配置中看到对应Pod相关的采集目标及详情,需要协助排查解决该问题。
排查修复方案
按以下优先级逐一校验修复:
- 补全PodMonitor命名空间选择器:当前Prometheus资源仅配置了
serviceMonitorNamespaceSelector筛选ServiceMonitor所在命名空间,缺失podMonitorNamespaceSelector配置。默认规则下Prometheus只会在自身部署的命名空间查找PodMonitor资源,跨命名空间部署的PodMonitor不会被识别。
若应用和PodMonitor部署在打了prometheus-scrape: "true"标签的命名空间,在Prometheus的spec段添加如下配置:
若需要允许所有命名空间的PodMonitor被发现,直接配置为podMonitorNamespaceSelector: matchLabels: prometheus-scrape: "true"podMonitorNamespaceSelector: {}即可。 - 核对全链路标签匹配:
- PodMonitor自身的
metadata.labels必须和Prometheus配置的podMonitorSelector规则匹配,当前PodMonitor带有app.kubernetes.io/name: prometheus-example-app标签,符合匹配要求,注意排查是否存在标签拼写错误、首尾多余空格的问题 - PodMonitor的
spec.selector.matchLabels必须和业务Pod的标签完全匹配,当前Deployment的Pod模板已打上对应标签,配置符合要求 - PodMonitor中配置的采集端口名
web必须和Pod容器定义的端口名完全一致,当前容器端口命名为web,配置正确
- PodMonitor自身的
- 校验RBAC权限:Prometheus绑定的
prometheusServiceAccount必须拥有业务所在命名空间下pod、pod/metrics资源的get、list、watch权限,权限不足会直接导致资源发现失败。可查看Prometheus Pod的运行日志,搜索forbidden关键字确认是否存在权限报错,按需补全Role/ClusterRole和对应绑定即可。 - 配置生效验证:
- 配置更新后打开Prometheus UI的Configuration页面,搜索
podMonitor/prometheus-example-app关键字,确认对应采集配置已经被正确加载 - 进入Service Discovery页面,查看对应PodMonitor的发现条目,确认Pod是否被正常识别、标签匹配无异常
- 若目标已被发现但采集失败,直接访问业务Pod的
http://<podIP>:8080/metrics接口,确认指标接口可以正常返回响应。
- 配置更新后打开Prometheus UI的Configuration页面,搜索
内容的提问来源于stack exchange,提问作者Dasari Sai Kumar
相关产品推荐
相关产品推荐

