OpenShift中Prometheus Blackbox Exporter自签名证书验证失败问题
问题描述
在已部署Prometheus Operator监控栈的OpenShift集群中,尝试用Blackbox Exporter探测Spring Boot应用的actuator/health端点,已完成以下操作:
- 在Prometheus Operator所属命名空间部署Blackbox Exporter,
Service与ConfigMap就绪,ConfigMap定义了http_2xx模块,Exporter运行正常 - 在两个命名空间部署相同应用,分别创建了静态目标的
Probe和动态发现的ServiceMonitor
当前探测均失败:
- ServiceMonitor报错:
level=info msg="Invalid HTTP response status code, wanted 2xx" status_code=400,添加scheme: https后仍无效 - Probe报错:
level=error msg="Error for HTTP request" err="Get \"https://appIP:port/actuator/health\": tls: failed to verify certificate: x509: certificate signed by unknown authority",尝试配置服务CA、应用证书密钥均无效,且tlsConfig: insecureSkipVerify: true未生效
相关配置
Blackbox Exporter ConfigMap
data: blackbox.yaml: | modules: http_2xx: http: no_follow_redirects: true method: GET preferred_ip_protocol: ip4 valid_http_versions: - HTTP/1.1 - HTTP/2 valid_status_codes: [] tls_config: insecure_skip_verify: true prober: http timeout: 10s
ServiceMonitor配置
spec: endpoints: - interval: 30s params: module: - http_2xx path: /probe relabelings: - action: replace sourceLabels: - __address__ targetLabel: __param_target - action: replace replacement: 'exporter:port' targetLabel: __address__ - action: replace sourceLabels: - __param_target targetLabel: instance - action: labelmap regex: __meta_kubernetes_service_label_(.+) scrapeTimeout: 10s jobLabel: jobLabel selector: matchLabels: app.kubernetes.io/component: component
Probe配置
spec: interval: 30s module: http_2xx prober: path: /probe url: 'exporter.namespace.svc:port' targets: staticConfig: static: - 'https://app.namespace.svc:port/actuator/health' tlsConfig: cert: secret: key: key name: secret-name keySecret: key: key name: secret-name
手动调用Blackbox Exporter日志
Logs for the probe: ts=2023-12-07T10:24:46.576847865Z caller=main.go:181 module=http_2xx target=https://app.namespace.svc:port level=info msg="Beginning probe" probe=http timeout_seconds=119.5 ts=2023-12-07T10:24:46.576945405Z caller=http.go:328 module=http_2xx target=https://app.namespace.svc:port level=info msg="Resolving target address" target=app.namespace.svc ip_protocol=ip4 ts=2023-12-07T10:24:46.615450737Z caller=http.go:328 module=http_2xx target=https://app.namespace.svc:port level=info msg="Resolved target address" target=app.namespace.svc ip=IP_of_service ts=2023-12-07T10:24:46.615543908Z caller=client.go:252 module=http_2xx target=https://app.namespace.svc:port level=info msg="Making HTTP request" url=https://IPaddress:port host=app.namespace.svc:port ts=2023-12-07T10:24:46.624148963Z caller=handler.go:120 module=http_2xx target=https://app.namespace.svc:port level=error msg="Error for HTTP request" err="Get \"https://IPaddress:port\": tls: failed to verify certificate: x509: certificate signed by unknown authority" ts=2023-12-07T10:24:46.624187979Z caller=handler.go:120 module=http_2xx target=https://app.namespace.svc:port level=info msg="Response timings for roundtrip" roundtrip=0 start=2023-12-07T10:24:46.618548821Z dnsDone=2023-12-07T10:24:46.618548821Z connectDone=2023-12-07T10:24:46.619955324Z gotConn=0001-01-01T00:00:00Z responseStart=0001-01-01T00:00:00Z tlsStart=2023-12-07T10:24:46.619998796Z tlsDone=2023-12-07T10:24:46.624134551Z end=0001-01-01T00:00:00Z ts=2023-12-07T10:24:46.62420857Z caller=main.go:181 module=http_2xx target=https://app.namespace.svc:port level=error msg="Probe failed" duration_seconds=0.047321017
解决方案
一、ServiceMonitor 400错误修复
当前ServiceMonitor的relabel规则仅传递了服务的ip:port作为探测目标,未拼接完整的HTTPS端点路径,且未考虑端点可能的认证要求:
- 修改relabel规则,拼接完整的HTTPS探测地址:
relabelings: - action: replace sourceLabels: - __address__ targetLabel: __param_target replacement: "https://$1/actuator/health" # 拼接完整端点URL - action: replace replacement: 'exporter:port' targetLabel: __address__ - action: replace sourceLabels: - __param_target targetLabel: instance - action: labelmap regex: __meta_kubernetes_service_label_(.+)
- 如果
actuator/health需要认证,在Blackbox的http_2xx模块中添加对应配置:
modules: http_2xx: http: # 保留原有配置... basic_auth: username: "your-username" password: "your-password" # 或使用Bearer Token bearer_token_file: "/var/run/secrets/tokens/your-service-account-token"
二、Probe证书验证失败修复
核心误区
Probe的tlsConfig是配置Prometheus与Blackbox Exporter之间的TLS通信,而非Blackbox与目标应用之间的TLS验证。目标应用的TLS配置需在Blackbox模块中定义,或通过Probe的moduleOverride覆盖。
具体修复方案
方案1:确保Blackbox模块的insecure_skip_verify生效
检查Blackbox Exporter Pod是否加载了最新的ConfigMap:
oc exec -it <blackbox-pod-name> -n <namespace> -- cat /etc/blackbox/blackbox.yaml
若配置未更新,重启Pod:
oc delete pod <blackbox-pod-name> -n <namespace>
方案2:使用OpenShift服务CA证书验证
若不想跳过证书验证,将OpenShift服务CA挂载到Blackbox Exporter并配置:
- 修改Blackbox Deployment,挂载服务CA:
volumes: - name: service-ca configMap: name: openshift-service-ca.crt containers: - name: blackbox-exporter volumeMounts: - name: service-ca mountPath: /var/run/secrets/kubernetes.io/serviceaccount readOnly: true
- 更新Blackbox模块的TLS配置:
tls_config: ca_file: "/var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt"
方案3:通过Probe的moduleOverride单独配置TLS
针对该Probe单独覆盖模块的TLS设置:
spec: interval: 30s module: http_2xx moduleOverride: http: tls_config: insecure_skip_verify: true prober: path: /probe url: 'exporter.namespace.svc:port' targets: staticConfig: static: - 'https://app.namespace.svc:port/actuator/health' # 移除原有的tlsConfig(仅用于Prometheus连接Blackbox,此处不需要)
内容的提问来源于stack exchange,提问作者Seidz_q
相关产品推荐
相关产品推荐

