You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OTEL Collector尾采样配置问题:Span Name过滤与采样规则调整

OpenTelemetry Collector Tail Sampling 配置解决方案

核心逻辑说明

tail_sampling 是基于完整Trace链路的采样机制,而非单个Span。要实现你的需求,必须从Trace维度制定规则:

  • 最高优先级:保留所有包含错误状态Span的Trace(无论是否属于MyEvent)
  • 次优先级:仅对MyEvent开头且全链路无错误的Trace,应用10%采样率(过滤90%)
  • 默认规则:所有未匹配上述规则的Trace(非MyEvent的所有请求)全量保留

最简有效配置

receivers:
  otlp:
    protocols:
      grpc:
      http:

processors:
  tail_sampling:
    decision_wait: 5s  # 确保收集完整Trace后再决策,可根据链路实际时长调整
    policies:
      # 策略1:保留所有含错误Span的Trace(优先级最高)
      - name: preserve-error-traces
        type: status_code
        status_code:
          status_codes: [ERROR]
        sampling_percentage: 100.0
      # 策略2:对MyEvent全链路成功的Trace采样10%
      - name: myevent-success-sample
        type: and
        and:
          and_sub_policy:
            # 匹配span.name为MyEvent的Trace
            - type: string_attribute
              string_attribute:
                key: "span.name"
                values: ["MyEvent"]
                match_type: exact
            # 排除含错误Span的Trace(已被策略1处理)
            - type: status_code
              status_code:
                status_codes: [ERROR]
                mode: exclude
        sampling_percentage: 10.0
      # 策略3:所有其他Trace全量保留
      - name: preserve-all-others
        type: always_sample

exporters:
  # 替换为你的实际存储Exporter(如Jaeger、OTLP等)
  jaeger:
    endpoint: "jaeger:14250"
    tls:
      insecure: true

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [tail_sampling]
      exporters: [jaeger]

关键配置修正点

  1. 错误链路全量保留:第一个策略优先执行,只要Trace中存在任何状态为ERROR的Span(包括子Span),就100%保留,彻底解决错误链路丢失问题。
  2. 精准匹配MyEvent成功链路:通过and组合两个子策略,仅对span.name为MyEvent且全链路无错误的Trace应用10%采样率,确保90%的成功请求被过滤。
  3. 默认全量保留:最后一个always_sample策略保证非MyEvent的所有Trace不受影响,完全遵循原有采样规则。
  4. decision_wait参数:设置足够时长等待Trace的所有Span上报,避免因子Span延迟到达导致误判为成功链路。

解决你遇到的三个问题

  • 无法按Span Name过滤:通过string_attribute策略精准匹配span.name属性,match_type: exact确保只命中名称为MyEvent的Trace。
  • 无法保留错误子Span链路:最高优先级的status_code策略直接拦截所有含错误的Trace,不会进入后续采样逻辑。
  • 采样比例不符合要求:仅对MyEvent全链路成功的Trace应用10%采样,其余场景全量保留,严格符合过滤90%成功请求的需求。

内容的提问来源于stack exchange,提问作者James

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 22:29:55