如何为OpenTelemetry-Collector-Tempo架构搭建Grafana APM风格仪表盘?
解决方案:补充指标采集链路,结合Prometheus+Tempo搭建APM仪表盘
首先明确:Tempo 是分布式追踪存储,仅存储链路追踪数据;而你需要的请求数、状态码分布、请求耗时等属于指标数据,这类数据需要用 Prometheus(或其他指标存储)来存储。所以要搭建符合需求的APM仪表盘,需要补充指标采集链路,同时结合 Tempo 的追踪能力实现完整的APM体验。
一、调整架构:新增指标采集链路
当前架构是 Rails -> OpenTelemetry Collector -> Tempo,需要扩展为:Rails -> OpenTelemetry Collector -> [Tempo(追踪) + Prometheus(指标)]
二、配置Rails端:开启指标采集
你的现有gem已经包含基础依赖,需要在OpenTelemetry初始化配置中开启指标采集,并配置OTLP指标导出。
修改Rails的初始化文件(比如 config/initializers/opentelemetry.rb):
require 'opentelemetry/sdk' require 'opentelemetry/exporter/otlp' require 'opentelemetry/instrumentation/rails' require 'opentelemetry/instrumentation/mysql2' require 'opentelemetry/instrumentation/net_http' require 'opentelemetry/instrumentation/rack' OpenTelemetry::SDK.configure do |c| # 配置追踪导出到OTLP(Collector) c.trace_exporter = OpenTelemetry::Exporter::OTLP::Exporter.new( endpoint: 'http://your-collector-host:4318/v1/traces' ) # 新增:配置指标导出到OTLP(Collector) c.metric_exporter = OpenTelemetry::Exporter::OTLP::Exporter.new( endpoint: 'http://your-collector-host:4318/v1/metrics' ) # 启用所有已安装的instrumentation,包括指标采集 c.use_all() end
三、配置OpenTelemetry Collector:同时转发追踪和指标
修改Collector的配置文件(比如 otel-collector-config.yaml),添加Prometheus exporter,将指标转发到Prometheus:
receivers: otlp: protocols: http: endpoint: "0.0.0.0:4318" grpc: endpoint: "0.0.0.0:4317" processors: batch: exporters: # 保持追踪导出到Tempo otlp/tempo: endpoint: "http://your-tempo-host:4317" tls: insecure: true # 新增:指标导出到Prometheus(支持remote write) prometheus: endpoint: "http://your-prometheus-host:9090/api/v1/write" tls: insecure: true service: pipelines: traces: receivers: [otlp] processors: [batch] exporters: [otlp/tempo] metrics: receivers: [otlp] processors: [batch] exporters: [prometheus]
四、在Grafana中搭建APM仪表盘
- 先配置Prometheus数据源:在Grafana中添加Prometheus数据源,指向你的Prometheus实例。
- 核心面板配置示例:
1. 总请求数(按状态码分组)
- 数据源:Prometheus
- 查询语句:
sum by (status_code) (rate(http_requests_total{service="your-rails-service-name"}[5m])) - 可视化类型:柱状图,展示各状态码的请求QPS
2. 请求耗时分位数(P50/P95/P99)
- 数据源:Prometheus
- 查询语句:
histogram_quantile(0.50, sum by (le) (rate(http_request_duration_seconds_bucket{service="your-rails-service-name"}[5m]))) histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket{service="your-rails-service-name"}[5m]))) histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{service="your-rails-service-name"}[5m]))) - 可视化类型:折线图,展示不同分位数的耗时趋势
3. 错误请求率(4xx/5xx占比)
- 数据源:Prometheus
- 查询语句:
sum(rate(http_requests_total{service="your-rails-service-name", status_code=~"4..|5.."}[5m])) / sum(rate(http_requests_total{service="your-rails-service-name"}[5m])) - 可视化类型:单数值面板,展示错误率百分比
4. 追踪关联(从指标跳转到Tempo)
在上述指标面板的面板链接中配置:
- 类型:Tempo
- 查询:
service.name="your-rails-service-name" http.status_code="{{status_code}}"(根据面板维度调整) - 时间范围:使用当前面板时间范围
这样点击指标面板的数据点时,就能直接跳转到Tempo中查看对应的链路追踪。
五、可选:复用现有仪表盘调整
你之前尝试的ID:19419仪表盘是基于Prometheus的,现在配置好Prometheus数据源后,可以导入该仪表盘,然后添加Tempo的追踪关联面板,就能快速得到完整的APM仪表盘。
内容的提问来源于stack exchange,提问作者fguillen
相关产品推荐
相关产品推荐

