OpenTelemetry Python中Trace Span起止点及插桩技术咨询
OpenTelemetry Python 追踪 Span 边界与插桩实战解析
1. 入站HTTP请求的Span起止位置
以Flask/Django这类Web框架的自动插桩为例:
- 创建点:请求进入框架的入口环节,比如Flask是在
wsgi_app被调用前(通过中间件包装实现),Django则是在请求中间件的process_request阶段附近。 - 结束点:请求处理完成、响应准备返回客户端的时刻,比如Flask的
wsgi_app执行完毕后,不管请求成功还是抛出异常都会结束Span。
拿opentelemetry-instrumentation-flask的源码来看,核心逻辑在FlaskInstrumentor._instrument方法里,它会替换原始的Flask.wsgi_app,在调用原始方法前创建Span,执行完后调用span.end()。
手动模拟入站Span的代码示例:
from opentelemetry import trace from flask import Flask app = Flask(__name__) tracer = trace.get_tracer(__name__) @app.route("/") def hello(): # 对应自动插桩的入口创建点 with tracer.start_as_current_span("http.request") as span: span.set_attribute("http.method", "GET") span.set_attribute("http.path", "/") # 业务逻辑执行 return "Hello World" # with块结束自动调用span.end(),对应自动插桩的收尾
2. 出站请求的Span起止点
HTTP客户端(requests库)
自动插桩的opentelemetry-instrumentation-requests中:
- 创建点:调用
requests.get/post等方法后,底层PreparedRequest.send执行前。 - 结束点:收到响应(或请求抛出异常)后,响应处理完成的时刻。
源码在requests/instrumentation.py的_send包装函数里,先创建Span并设置http.url、http.method等标准属性,再调用原始send方法,最后结束Span。
数据库查询(psycopg2)
opentelemetry-instrumentation-psycopg2的Span:
- 创建点:
cursor.execute方法执行前。 - 结束点:查询完成(拿到结果集)或抛出异常后。
手动模拟出站HTTP请求的Span:
import requests from opentelemetry import trace tracer = trace.get_tracer(__name__) def fetch_external_data(): # 请求发起前创建Span with tracer.start_as_current_span("external.http.fetch") as span: span.set_attribute("http.url", "https://api.example.com/data") span.set_attribute("http.method", "GET") response = requests.get("https://api.example.com/data") # 响应接收后设置状态码,with块结束自动结束Span span.set_attribute("http.status_code", response.status_code) return response.json()
3. 内部本地函数的Span边界与插桩
自动插桩一般通过装饰器注入或字节码修改实现,Span起始于函数调用瞬间,结束于函数返回/抛出异常时刻。手动插桩有两种常用方式:
方式1:装饰器
from opentelemetry import trace tracer = trace.get_tracer(__name__) @tracer.start_as_current_span("user.profile.process") def process_user_profile(user_id): # 核心业务逻辑 return {"user_id": user_id, "status": "processed"}
方式2:上下文管理器
def calculate_order_stats(): with tracer.start_as_current_span("order.stats.calculate") as span: # 计算逻辑 span.set_attribute("stats.total_orders", 150) return {"total": 150, "avg_value": 23.5}
边界确定原则:
- 优先给核心业务函数、性能敏感的计算步骤加Span,别给琐碎的工具函数插桩(会导致追踪数据冗余)
- Span命名要贴合业务含义,比如
user.profile.load比load_data更清晰
自动vs手动插桩对比方案
如果看源码没搞懂自动插桩逻辑,直接手动插桩对比是最直观的:
- 先开启自动插桩,跑一次请求,导出追踪数据(用Jaeger/OTel Collector),记录Span的起止时间、属性、嵌套关系
- 禁用自动插桩,用手动方式给相同的请求、出站调用、内部函数加Span,再导出数据
- 对比两者的Span结构:自动插桩会自动填充标准属性(如
http.status_code、db.statement),嵌套关系应该和手动实现一致
关键源码位置
- 通用Instrumentation框架:
opentelemetry-python仓库的opentelemetry/instrumentation目录,核心是instrumentor.py和base_instrumentor.py - Flask入站插桩:
opentelemetry-instrumentation-flask的src/opentelemetry/instrumentation/flask/__init__.py - Requests出站插桩:
opentelemetry-instrumentation-requests的src/opentelemetry/instrumentation/requests/instrumentation.py - Psycopg2数据库插桩:
opentelemetry-instrumentation-psycopg2的src/opentelemetry/instrumentation/psycopg2/__init__.py
内容的提问来源于stack exchange,提问作者Knut
相关产品推荐
相关产品推荐

