SpringBoot+OpenTelemetry+Micrometer环境下自定义Span适用场景咨询
在Spring Boot + OpenTelemetry + Micrometer的可观测性方案中,自动Instrumentation(如Spring MVC、JDBC、Redis等的默认追踪)能覆盖大部分通用链路,但以下场景下手动创建自定义Span是必要的:
1. 业务逻辑核心环节的拆分追踪
当你需要将复杂业务流程拆解为可独立观测的子步骤时,自定义Span能帮你精准定位每个业务节点的耗时、成功率。比如电商下单流程,自动追踪只会记录整个接口的Span,但拆分出「创建订单」「扣减库存」「发起支付」等子Span后,能直观看到哪个环节拖慢了整体流程。
示例代码:
@Autowired private Tracer tracer; public void submitOrder(OrderRequest request) { // 自动生成的根Span来自Spring MVC请求 try (Span createSpan = tracer.spanBuilder("business.order.create").startSpan()) { createSpan.setAttribute("order.id", request.getOrderId()); orderRepository.saveOrder(request); // 嵌套子Span:扣减库存 try (Span inventorySpan = tracer.spanBuilder("business.inventory.deduct").startSpan()) { inventorySpan.setAttribute("product.id", request.getProductId()); inventoryService.deduct(request.getProductId(), request.getQuantity()); } // 嵌套子Span:发起支付 try (Span paymentSpan = tracer.spanBuilder("business.payment.init").startSpan()) { paymentSpan.setAttribute("amount", request.getTotalAmount()); paymentGateway.initPayment(request); } } }
2. 未被自动Instrumentation覆盖的组件调用
OpenTelemetry的自动代理无法覆盖所有自定义或私有组件,比如:
- 自定义文件IO操作、本地缓存(如Caffeine)
- 未提供OpenTelemetry支持的第三方私有SDK
- 自研的内部服务调用协议
这类场景下,手动添加Span能填补链路空白,确保观测数据的完整性。
示例代码:
public void writeLocalCache(String key, Object value) { try (Span cacheSpan = tracer.spanBuilder("cache.local.put").startSpan()) { cacheSpan.setAttribute("cache.key", key); localCache.put(key, value); cacheSpan.setStatus(StatusCode.OK); } catch (Exception e) { cacheSpan.setStatus(StatusCode.ERROR, e.getMessage()); cacheSpan.recordException(e); throw new CacheOperationException("Failed to write cache", e); } }
3. 异步任务与消息队列的链路追踪
异步场景(如@Async方法、RabbitMQ/Kafka消费者)中,自动追踪可能无法维持完整的链路上下文——因为异步线程与请求线程是隔离的。此时需要手动创建Span并关联父Span的上下文,确保链路的连续性。
示例代码(@Async任务):
@Async public void processNotification(String userId, String message) { // 获取父Span上下文(来自请求线程) SpanContext parentContext = tracer.getCurrentSpan().getSpanContext(); try (Span notifySpan = tracer.spanBuilder("async.notification.send") .setParent(parentContext) .startSpan()) { notifySpan.setAttribute("user.id", userId); notificationService.send(userId, message); } }
4. 耗时较长的批处理/计算任务
对于大数据批量处理、报表生成、ETL等耗时任务,单个Span无法体现内部执行细节。拆分多个子Span能帮你定位到具体哪个子任务(如某条数据处理、某个计算步骤)出现延迟或失败。
示例代码:
public void batchExportUserData(List<String> userIds) { try (Span batchSpan = tracer.spanBuilder("batch.user.export").startSpan()) { batchSpan.setAttribute("total.count", userIds.size()); for (String userId : userIds) { try (Span exportSpan = tracer.spanBuilder("batch.user.export.single").startSpan()) { exportSpan.setAttribute("user.id", userId); userDataExporter.export(userId); } } } }
5. 业务异常与错误的精细化追踪
当需要记录业务级异常的上下文信息时,自定义Span可以附加详细的属性(如用户ID、资源ID)并标记错误状态,方便后续通过可观测平台快速筛选和排查问题。
示例代码:
public void validateUserAccess(String userId, String resource) { try (Span checkSpan = tracer.spanBuilder("auth.permission.check").startSpan()) { checkSpan.setAttribute("user.id", userId); checkSpan.setAttribute("resource.path", resource); if (!authService.hasPermission(userId, resource)) { checkSpan.setStatus(StatusCode.ERROR, "Permission denied"); checkSpan.recordException(new AccessDeniedException("User " + userId + " cannot access " + resource)); throw new AccessDeniedException("No permission to access resource"); } } }
核心原则
自定义Span的添加要遵循「必要且有价值」的原则:只在能帮你解决链路断点、性能瓶颈、故障排查问题的节点添加,避免过度创建导致链路冗余、观测成本上升。
内容的提问来源于stack exchange,提问作者HyoD.k

