You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenTracing最佳实践:应追踪内部操作还是仅API调用?

OpenTracing 追踪粒度与范围的最佳实践

Great questions—these are some of the most common pitfalls when rolling out distributed tracing, so let’s break them down with practical context.

问题一:应该追踪内部服务工作,还是仅追踪API调用?

There’s no one-size-fits-all answer here—it depends entirely on what you’re trying to achieve with tracing:

  • If your top priority is cross-service request troubleshooting: Stick to tracking API calls (inbound/outbound requests between services) first. This gives you the high-level map of how a request flows through your system, which is critical for pinpointing which service is causing delays or failures in a distributed chain.
  • If you need to diagnose internal performance bottlenecks: You’ll want to add tracing for key internal workflows. For example, if an API endpoint is slow, knowing whether the delay comes from a complex internal calculation, a slow database query, or a blocked thread requires visibility into those internal steps.
  • A balanced starting point: Begin with coarse-grained tracing (cross-service APIs) to establish your core request map. Then, incrementally add spans for internal operations only when you’ve identified a need—like repeatedly troubleshooting slowdowns in a specific service component. Avoid over-tracing every tiny internal step upfront; it adds unnecessary overhead and creates noise in your tracing data.

问题二:通知处理流程中的span粒度选择

For your specific workflow (receive input → process [calculations, DB storage, generate notifications] → send to clients), here’s how to balance tracing spans and metrics tools like Prometheus:

1. Start with a top-level span

First, create a single top-level span for the entire "input notification processing" workflow. This spans from the moment your service receives the input notification to when all client notifications are sent. It gives you the big picture of how long the entire job takes, and acts as the parent for any child spans you add.

2. Add child spans for critical, high-impact steps

You’ll want separate spans for operations that meet any of these criteria:

  • External dependencies: Database storage operations are a perfect example. While Prometheus can track average DB latency, a span ties that latency to a specific request context (e.g., "this DB write took 500ms while processing notification X"). This is invaluable when debugging why a single request failed or was slow.
  • Long-running or resource-heavy operations: If your calculation step involves complex logic (e.g., data transformation, machine inference) that takes more than a few hundred milliseconds, a child span will help you isolate whether this step is the bottleneck.
  • Core business logic: The "generate own notification" step is a key business action—adding a span here makes it easier to trace how this business logic executes across requests, and flag if it’s failing or behaving unexpectedly.
  • Multi-client notification sending: If sending to each client involves separate external calls (e.g., HTTP requests to client APIs), consider a child span per client, or a single span for the entire "send notifications" batch (whichever gives you clearer visibility into failures/delays).

3. Leave small, high-frequency operations to Prometheus

For simple, short-lived internal calculations (e.g., basic arithmetic, data formatting) that take milliseconds to complete, don’t add individual spans. Instead, use Prometheus to track metrics like:

  • Average time spent on calculation steps per request
  • Total number of calculation operations executed
  • Error rates for calculation logic

Metrics tools are better suited for aggregating this kind of data, and adding spans here would only bloat your tracing system without providing meaningful troubleshooting value.

Final breakdown for your workflow

  • Top-level span: Process Input Notification (start: receive input, end: all client notifications sent)
  • Child spans:
    • Store Data to Database
    • Generate Internal Notification
    • Send Notifications to Clients (or individual spans per client if needed)
  • Metrics-only operations: Simple, fast calculation steps (track via Prometheus gauges/counters)

Key Takeaway

The golden rule for tracing is: only add spans that solve a specific problem (troubleshooting a slowdown, debugging a failure, verifying a business workflow). If a span doesn’t give you actionable insight, it’s just overhead. Balance is everything—tracing and metrics tools work best together, not in competition.

内容的提问来源于stack exchange,提问作者Michał Szewczyk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:24:23