You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenTelemetry指标缺失K8s属性,仅显示自身Pod信息求助

问题描述

在EKS集群中配置OpenTelemetry后,Traces能正常携带K8s属性,可在Grafana结合Tempo查看,但Metrics仅包含OTel自身Pod和Namespace信息,缺少业务应用的Pod名称、Namespace等K8s相关属性,目前只能通过service_name标签识别微服务,希望指标中能带上这些K8s详情。

示例指标:

calls_total{container="otelcol", db_operation="DELETE", db_statement="DELETE FROM epod001_arosaas_test_sig.QRTZ_FIRED_TRIGGERS WHERE SCHED_NAME = ? AND ENTRY_ID = ?", endpoint="metrics", instance="10.220.7.235:6666", job="opentelemetrycollector", namespace="monitoring", operation="DELETE epod001_arosaas_test_sig.QRTZ_FIRED_TRIGGERS", pod="opentelemetrycollector-5d5885d69d-tb8s9", service="opentelemetrycollector", service_name="aro-batch-jobs", span_kind="SPAN_KIND_CLIENT", status_code="STATUS_CODE_UNSET"}
现有配置

OpenTelemetry Collector配置(ConfigMap)

apiVersion: v1
kind: ConfigMap
metadata:
  name: collector-config
  namespace: monitoring
  labels:
    app: opentelemetry
    component: otel-collector-conf
data:
  collector.yaml: |
    receivers:
      # Make sure to add the otlp receiver.
      # This will open up the receiver on port 4317
      otlp:
        protocols:
         grpc:
           endpoint: "0.0.0.0:5555"
         http:
      hostmetrics:
        collection_interval: 20s
        scrapers:
          cpu:
            metrics:
             system.cpu.utilization:
               enabled: true
          load:
          memory:
            metrics:
              system.memory.utilization:
                enabled: true
          disk:
          filesystem:
            metrics:
              system.filesystem.utilization:
                enabled: true
          network:
          paging:
            metrics:
              system.paging.utilization:
                enabled: true
          processes:
          process:
      k8s_cluster:
        collection_interval: 10s
        node_conditions_to_report: [Ready, MemoryPressure,DiskPressure,NetworkUnavailable]
        allocatable_types_to_report: [cpu, memory,storage]
      k8s_events:
        auth_type : serviceAccount
      receiver_creator:
         watch_observers: [k8s_observer]
         receivers:
            kubeletstats:
              rule: type == "k8s.node"
              config:
                collection_interval: 10s
                auth_type: serviceAccount
                endpoint: "`endpoint`:`kubelet_endpoint_port`"
                insecure_skip_verify: true
                extra_metadata_labels:
                  - container.id
                  - k8s.volume.type
                metric_groups:
                  - node
                  - pod
                  - volume
                  - container
      prometheus:
        config:
          scrape_configs:
            - job_name: 'kube-state-metrics'
              scrape_interval: 5s
              scrape_timeout: 1s
              static_configs:
                - targets: ['kube-state-metrics.monitoring.svc.cluster.local:8443']
            - job_name: k8s
              kubernetes_sd_configs:
              - role: pod
              relabel_configs:
              - source_labels: ['__meta_kubernetes_pod_annotation_prometheus_io_scrape']
                regex: "true"
                action: keep
              metric_relabel_configs:
              - source_labels: [__name__]
                regex: "(request_duration_seconds.*|response_duration_seconds.*)"
                action: keep
    processors:
      memory_limiter:
        check_interval: 1s
        limit_mib: 2000
        spike_limit_mib: 500
      batch:
        timeout: 10s
        send_batch_size: 10000
      spanmetrics:
        metrics_exporter: prometheus
        latency_histogram_buckets: [100ms, 250ms,500ms,1s,2s,4s,6s,8s,10s,20s,30s]
        dimensions:
          - name: http.method
          - name: http.status_code
          - name: db.operation
          - name: db.statement
          - name: exception.message
          - name: exception.type
          - name: messaging.message.id
          - name: messaging.message.payload_size_bytes
        dimensions_cache_size: 10000
        aggregation_temporality: "AGGREGATION_TEMPORALITY_CUMULATIVE"
      servicegraph:
        metrics_exporter: prometheus
      transform:
         metric_statements:
           - context: metric
             statements:
               - set(description, "Measures the duration of inbound HTTP requests") where name == "http.server.duration"
      cumulativetodelta:
        include:
          metrics:
           - system.network.io
           - system.disk.operations
           - system.network.dropped
           - system.network.packets
           - process.cpu.time
          match_type: strict
      resource/k8s:
        attributes:
          - key: host.id
            from_attribute: host.name
            action: upsert
          - key: k8s.cluster.name
            from_attribute: sig-test-ekscluster-001
            action: insert
      resourcedetection:
        detectors: [env, system]
      k8sattributes:
        auth_type: serviceAccount
        passthrough: false
        filter:
          node_from_env_var: K8S_NODE_NAME
        extract:
          metadata:
            - k8s.pod.name
            - k8s.pod.uid
            - k8s.deployment.name
            - k8s.namespace.name
            - k8s.node.name
            - k8s.pod.start_time
      metricstransform:
        transforms:
           include: .+
           match_type: regexp
           action: update
           operations:
             - action: add_label
               new_label: kubernetes.cluster.id
               new_value: CLUSTER_ID_TO_REPLACE
             - action: add_label
               new_label: kubernetes.name
               new_value: ezops-client-test-saas-eks-001
    extensions:
      health_check: {}
    exporters:
      otlp:
        endpoint: "http://tempo.monitoring.svc.cluster.local:55680"
        tls:
          insecure: true
      prometheus:
        endpoint: "0.0.0.0:6666"
      logging:
        loglevel: debug
      loki:
         endpoint: http://loki.monitoring.svc.cluster.local:3100/loki/api/v1/push
         labels:
           resource:
             container.name: "container_name"
             k8s.cluster.name: "k8s_cluster_name"
             k8s.event.reason: "k8s_event_reason"
             k8s.object.kind: "k8s_object_kind"
             k8s.object.name: "k8s_object_name"
             k8s.object.uid: "k8s_object_uid"
             k8s.object.fieldpath: "k8s_object_fieldpath"
             k8s.object.api_version: "k8s_object_api_version"
           attributes:
             k8s.event.reason: "k8s_event_reason"
             k8s.event.action: "k8s_event_action"
             k8s.event.start_time: "k8s_event_start_time"
             k8s.event.name: "k8s_event_name"
             k8s.event.uid: "k8s_event_uid"
             k8s.namespace.name: "k8s_namespace_name"
             k8s.event.count: "k8s_event_count"
           record:
             traceID: "traceid"
    service:
      extensions: [health_check]
      pipelines:
          logs:
            receivers: [k8s_events]
            processors: [memory_limiter,k8sattributes,batch]
            exporters: [loki,logging]
          traces:
            receivers: [otlp]
            processors: [spanmetrics,servicegraph,batch]
            exporters: [otlp]
          metrics:
            receivers: [otlp,prometheus]
            processors: [memory_limiter,metricstransform,k8sattributes,resourcedetection,batch,resource/k8s]
            exporters: [logging,prometheus,otlp]
      telemetry:
        logs:
          level: debug
          initial_fields:
            service: my-prom-instance

RBAC配置

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: otelcontribcol
  labels:
    app: otelcontribcol
rules:
  - apiGroups:
      - ""
    resources:
      - events
      - namespaces
      - namespaces/status
      - nodes
      - nodes/spec
      - nodes/stats
      - nodes/proxy
      - pods
      - pods/status
      - replicationcontrollers
      - replicationcontrollers/status
      - resourcequotas
      - services
    verbs:
      - get
      - list
      - watch
  - apiGroups:
      - apps
    resources:
      - daemonsets
      - deployments
      - replicasets
      - statefulsets
    verbs:
      - get
      - list
      - watch
  - apiGroups:
      - extensions
    resources:
      - daemonsets
      - deployments
      - replicasets
    verbs:
      - get
      - list
      - watch
  - apiGroups:
      - batch
    resources:
      - jobs
      - cronjobs
    verbs:
      - get
      - list
      - watch
  - apiGroups:
      - autoscaling
    resources:
      - horizontalpodautoscalers
    verbs:
      - get
      - list
      - watch
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: otelcontribcol
  labels:
    app: otelcontribcol
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: otelcontribcol
subjects:
  - kind: ServiceAccount
    name: prometheus-k8s
    namespace: monitoring
解决方案

1. 调整Metrics Pipeline中处理器顺序

当前Metrics Pipeline中k8sattributes在metricstransform之后执行,会导致K8s属性无法被后续标签转换正确处理。将k8sattributes移到metricstransform之前:

metrics:
  receivers: [otlp,prometheus]
  processors: [memory_limiter,k8sattributes,metricstransform,resourcedetection,batch,resource/k8s]
  exporters: [logging,prometheus,otlp]

2. 完善k8sattributes的Pod关联规则

默认关联规则依赖k8s.pod.uid,但部分业务指标可能未携带该属性。添加基于service.name的关联规则,匹配对应的Deployment:

k8sattributes:
  auth_type: serviceAccount
  passthrough: false
  filter:
    node_from_env_var: K8S_NODE_NAME
  pod_association:
    - sources:
        - from: resource_attribute
          name: k8s.pod.uid
    - sources:
        - from: resource_attribute
          name: service.name
          regex: "^(.*)$"
          target: k8s.deployment.name
  extract:
    metadata:
      - k8s.pod.name
      - k8s.pod.uid
      - k8s.deployment.name
      - k8s.namespace.name
      - k8s.node.name
      - k8s.pod.start_time

3. 增强Prometheus Receiver的Relabel规则

在Prometheus的k8s抓取任务中,直接将K8s元数据注入为指标标签,确保从业务Pod抓取的指标自带基础K8s属性:

- job_name: k8s
  kubernetes_sd_configs:
  - role: pod
  relabel_configs:
  - source_labels: ['__meta_kubernetes_pod_annotation_prometheus_io_scrape']
    regex: "true"
    action: keep
  - source_labels: ['__meta_kubernetes_pod_name']
    target_label: 'k8s.pod.name'
  - source_labels: ['__meta_kubernetes_namespace']
    target_label: 'k8s.namespace.name'
  - source_labels: ['__meta_kubernetes_pod_uid']
    target_label: 'k8s.pod.uid'
  metric_relabel_configs:
  - source_labels: [__name__]
    regex: "(request_duration_seconds.*|response_duration_seconds.*)"
    action: keep

4. 验证RBAC权限

当前RBAC已包含pods、namespaces等资源的get/list/watch权限,prometheus-k8s ServiceAccount可正常访问相关元数据,无需额外调整。

验证

修改配置后重启OpenTelemetry Collector Pod,查看生成的指标是否包含k8s.pod.name、k8s.namespace.name等标签,示例如下:

calls_total{container="aro-batch-jobs", db_operation="DELETE", ..., k8s.pod.name="aro-batch-jobs-xyz-123", k8s.namespace.name="business", service_name="aro-batch-jobs", ...}

内容的提问来源于stack exchange,提问作者Arpit Tyagi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 20:32:32