You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

启用Tail Sampling后OpenTelemetry Collector高负载下停止导出数据求助

问题描述

有多台通过OpenTelemetry Agent 1.21.0埋点的JVM,仅将链路追踪数据发送至由Nginx代理的5台OpenTelemetry Collector 0.70.0(网关部署模式),后端对接Grafana Tempo。遇到的问题是:当JVM服务RPS上升时,Collector会停止导出数据——Tail Sampling配置在RPS低于1K时运行正常,但负载超过该值后无法导出;关闭Tail Sampling处理器后,即使RPS高达150K也能正常工作。

当前Collector配置
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317

processors:
  tail_sampling:
    decision_wait: 3s
    num_traces: 100
    expected_new_traces_per_sec: 10
    policies:
      - name: latency500-policy
        type: latency
        latency:
          threshold_ms: 500
      - name: error-policy
        type: string_attribute
        string_attribute:
          key: error
          values:
            - true
            - True
      - name: probabilistic10-policy
        type: probabilistic
        probabilistic:
          sampling_percentage: 10

exporters:
  otlp:
    endpoint: tempo:4317
    tls:
      insecure: true

service:
  extensions:
    - health_check
  pipelines:
    traces:
      receivers:
        - otlp
      processors:
        - tail_sampling
      exporters:
        - otlp

extensions:
  health_check:
    endpoint: 0.0.0.0:14319
    path: "/health"
问题根源与修复方案

核心问题:Tail Sampling配置参数严重不匹配实际负载

你的Tail Sampling处理器配置里,expected_new_traces_per_sec和num_traces这两个核心参数完全不符合高负载场景:

  • expected_new_traces_per_sec设为10,远低于实际峰值RPS(150K),导致Collector无法为高负载分配足够的缓存资源。
  • num_traces设为100,限制了内存中同时处理的追踪数上限,高RPS下会迅速填满缓存,后续追踪直接被丢弃,最终停止导出数据。

具体修复步骤

  1. 调整缓存核心参数

    • 将expected_new_traces_per_sec设置为接近实际峰值RPS的数值,比如100000(留足余量),让Collector能预判负载并分配对应资源。
    • 大幅提升num_traces的值,比如10000或更高,确保内存能容纳高负载下的待处理追踪。注意需结合Collector的内存资源调整,避免内存溢出。
  2. 优化决策等待时间
    当前decision_wait为3秒,高负载下可适当缩短至1-2秒,减少追踪在内存中的停留时间,降低缓存压力。

  3. 补充Collector资源配额
    Tail Sampling属于CPU和内存密集型操作,高负载下需为5台Collector分配足够的CPU和内存资源,避免资源瓶颈导致处理中断。

  4. 升级Collector版本
    你使用的0.70.0版本较旧,后续版本对Tail Sampling的性能和稳定性有大量优化,建议升级到最新稳定版(如v0.90+),能有效解决旧版本的高负载适配问题。

调整后的参考配置示例

processors:
  tail_sampling:
    decision_wait: 2s
    num_traces: 10000
    expected_new_traces_per_sec: 100000
    policies:
      - name: latency500-policy
        type: latency
        latency:
          threshold_ms: 500
      - name: error-policy
        type: string_attribute
        string_attribute:
          key: error
          values:
            - true
            - True
      - name: probabilistic10-policy
        type: probabilistic
        probabilistic:
          sampling_percentage: 10

内容的提问来源于stack exchange,提问作者Dmitrii Apanasevich

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 15:57:38