You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenTelemetry Target Allocator未分配已发现的ServiceMonitors求助

问题分析与解决方案

问题核心在于两点:一是手动配置的scrape_configs导致TargetAllocator(TA)仅关联自定义空job;二是TA默认未配置为扫描所有Namespace的ServiceMonitor,且权限可能不足。

解决步骤

1. 移除手动定义的Prometheus采集配置

当启用targetAllocator.prometheusCR.enabled: true时,Operator会自动生成基于ServiceMonitor的服务发现配置,无需手动编写scrape_configs。保留手动配置会导致TA忽略其他ServiceMonitor生成的job。

2. 配置TargetAllocator扫描所有Namespace的ServiceMonitor

在targetAllocator.prometheusCR段添加选择器,确保TA能发现所有Namespace下的ServiceMonitor:

  • namespaceSelector: {} 表示匹配所有Namespace
  • selector: {} 表示匹配所有ServiceMonitor(如需过滤,可添加标签规则,例如matchLabels: {app: metrics})

3. 赋予ServiceAccount跨Namespace权限

TA需要读取所有Namespace的ServiceMonitor、Pod、Service等资源,需创建对应的ClusterRole和ClusterRoleBinding:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: opentelemetry-targetallocator-clusterrole
rules:
- apiGroups: ["monitoring.coreos.com"]
  resources: ["servicemonitors"]
  verbs: ["get", "list", "watch"]
- apiGroups: [""]
  resources: ["services", "pods", "endpoints"]
  verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: opentelemetry-targetallocator-clusterrolebinding
subjects:
- kind: ServiceAccount
  name: opentelemetry-targetallocator-sa
  namespace: otel-metrics
roleRef:
  kind: ClusterRole
  name: opentelemetry-targetallocator-clusterrole
  apiGroup: rbac.authorization.k8s.io

4. 修改后的OpenTelemetryCollector配置

apiVersion: opentelemetry.io/v1alpha1
kind: OpenTelemetryCollector
metadata:
  name: otel-metrics
  namespace: otel-metrics
spec:
  mode: statefulset
  targetAllocator:
    enabled: true
    serviceAccount: opentelemetry-targetallocator-sa
    prometheusCR:
      enabled: true
      namespaceSelector: {}
      selector: {}
  config: |
    receivers:
      prometheus:
        config: {}

    exporters:
      logging:
        verbosity: detailed
      prometheus:
        endpoint: "0.0.0.0:8889"
        send_timestamps: true
        metric_expiration: 180m

    service:
      pipelines:
        metrics:
          receivers:
          - prometheus
          processors: []
          exporters:
          - logging
          - prometheus
      telemetry:
        logs:
          level: "debug"

验证配置

部署修改后的资源后,查看Collector的ConfigMap,会发现Prometheus Receiver的配置被自动替换为全局服务发现URL:

receivers:
  prometheus:
    config:
      scrape_configs:
      - http_sd_configs:
        - url: http://otel-metrics-targetallocator:80/jobs/all/targets?collector_id=$POD_NAME
          refresh_interval: 1m
        job_name: otel-metrics-all

此时访问该URL即可获取所有ServiceMonitor生成的采集目标。

内容的提问来源于stack exchange,提问作者Stuart Buckingham

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 15:10:24