You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

K8s中Otel Collector无集群级Pod列表权限问题求助

问题根源分析

错误提示pods is forbidden: User "system:serviceaccount:monitoring:otel-collector-collector" cannot list resource "pods" in API group "" at the cluster scope,核心原因有两点:

  1. ServiceAccount不匹配:你部署的OpenTelemetryCollector CR会自动生成名为otel-collector-collector的ServiceAccount,而你手动创建的otel-collector-sa并未被Collector实例指定使用,所以Pod实际使用的是自动生成的SA。
  2. 权限未绑定:你现有的prometheus ClusterRole虽然包含了pods的list/watch权限,但对应的ClusterRoleBinding只把这个角色绑定给了monitoring命名空间下的default ServiceAccount,没有绑定到Collector实际使用的SA上。
解决方案

可以选择以下两种方案中的任意一种:

方案一:给自动生成的SA绑定权限

创建一个新的ClusterRoleBinding,将prometheus ClusterRole绑定到otel-collector-collector SA:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: otel-collector-clusterrolebinding
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: prometheus
subjects:
- kind: ServiceAccount
  name: otel-collector-collector
  namespace: monitoring

执行命令应用配置:

kubectl apply -f <上述文件路径>

方案二:指定Collector使用自定义SA并绑定权限

  1. 修改你的otel-deployment.yaml,在spec中添加serviceAccount字段,指定使用你创建的otel-collector-sa:
apiVersion: opentelemetry.io/v1alpha1
kind: OpenTelemetryCollector
metadata:
  name: otel-collector
  namespace: monitoring
spec:
  serviceAccount: otel-collector-sa  # 新增该行,指定自定义SA
  config: |
    # 原有配置内容保持不变
    receivers:
      prometheus:
        config:
          scrape_configs:
            - job_name: 'kube-state-metrics'
              scrape_interval: 5s
              scrape_timeout: 1s
              static_configs:
                - targets: ['kube-state-metrics.kube-system.svc.cluster.local:8080']
            - job_name: k8s
              kubernetes_sd_configs:
              - role: pod
              relabel_configs:
              - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
                regex: "true"
                action: keep
              metric_relabel_configs:
              - source_labels: [__name__]
                regex: "(request_duration_seconds.*|response_duration_seconds.*)"
                action: keep
    processors:
      batch:
    exporters:
      logging:
    service:
      pipelines:
        metrics:
          receivers: [prometheus]
          exporters: [logging]
      telemetry:
        logs:
          level: debug
          initial_fields:
            service: my-prom-instance
  1. 创建ClusterRoleBinding,将prometheus ClusterRole绑定到otel-collector-sa:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: otel-collector-sa-clusterrolebinding
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: prometheus
subjects:
- kind: ServiceAccount
  name: otel-collector-sa
  namespace: monitoring
  1. 执行命令更新Collector配置和绑定权限:
kubectl apply -f otel-deployment.yaml
kubectl apply -f <上述ClusterRoleBinding文件路径>
验证

应用配置后,查看Collector Pod日志,确认权限错误消失:

kubectl logs -n monitoring <otel-collector-pod-name>

内容的提问来源于stack exchange,提问作者Ritesh Sinha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 12:06:44