You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何装饰Python进程以捕获其发起的所有HTTP请求与响应数据?

全量Python出站HTTP请求可编程捕获方案

下面两种方案均为进程内实现,可完全自定义请求/响应的处理逻辑,满足你提到的日志记录、事件上报、数据分析、测试集成需求。

方案1:基于OpenTelemetry埋点(推荐)

OpenTelemetry的Python生态已提供主流HTTP客户端的官方埋点实现,无需手动适配,不会漏抓请求,你可以自定义链路处理器完全在本地处理数据,不需要上报到任何外部服务。

实现步骤

  1. 安装对应依赖:
pip install opentelemetry-api opentelemetry-sdk opentelemetry-instrumentation-requests opentelemetry-instrumentation-urllib3 opentelemetry-instrumentation-aiohttp-client

你可以根据代码库用到的HTTP客户端按需安装对应instrumentation包,覆盖requests、urllib3、aiohttp、httpx、标准库urllib等几乎所有常用客户端。

  1. 自定义请求处理逻辑示例:
import logging
import pandas as pd
from opentelemetry import trace
from opentelemetry.sdk.trace import SpanProcessor
from opentelemetry.sdk.trace import TracerProvider

# 全局存储请求数据用于后续分析
global_http_req_df = pd.DataFrame(columns=[
    "method", "url", "status_code", "start_time", "end_time", "duration", "trace_id"
])

# 自定义链路处理器,所有HTTP请求的链路数据都会流入该处理器
class HttpReqCaptureProcessor(SpanProcessor):
    def on_start(self, span, parent_context):
        # 过滤非HTTP请求的链路
        if not span.attributes.get("http.method"):
            return
        # 暂存请求基础信息
        span.set_attribute("custom.req_captured", True)
    
    def on_end(self, span):
        if not span.attributes.get("custom.req_captured"):
            return
        # 组装完整请求响应数据
        req_data = {
            "method": span.attributes["http.method"],
            "url": span.attributes["http.url"],
            "status_code": span.attributes.get("http.status_code"),
            "start_time": span.start_time,
            "end_time": span.end_time,
            "duration": span.end_time - span.start_time,
            "trace_id": format(span.get_span_context().trace_id, "032x")
        }
        # 1. 用现有logger记录
        logging.info(f"HTTP请求捕获: {req_data['method']} {req_data['url']} 状态码: {req_data['status_code']}")
        # 2. 可在此处添加Kafka上报逻辑
        # 3. 追加到DataFrame用于后续分析
        global global_http_req_df
        global_http_req_df = pd.concat([global_http_req_df, pd.DataFrame([req_data])], ignore_index=True)

# 初始化链路追踪组件
trace_provider = TracerProvider()
trace_provider.add_span_processor(HttpReqCaptureProcessor())
trace.set_tracer_provider(trace_provider)

# 开启对应HTTP客户端的埋点,以requests为例
from opentelemetry.instrumentation.requests import RequestsInstrumentor
RequestsInstrumentor().instrument()

测试集成示例

在pytest/unittest中可以直接在测试用例的setup阶段初始化处理器,测试结束后直接断言捕获到的请求数据是否符合预期:

import pytest

@pytest.fixture(scope="function")
def captured_http_requests():
    req_list = []
    # 临时修改处理器逻辑,将请求存入req_list
    original_on_end = HttpReqCaptureProcessor.on_end
    def mock_on_end(self, span):
        if span.attributes.get("custom.req_captured"):
            req_list.append({
                "method": span.attributes["http.method"],
                "url": span.attributes["http.url"],
                "status_code": span.attributes.get("http.status_code")
            })
        return original_on_end(self, span)
    HttpReqCaptureProcessor.on_end = mock_on_end
    yield req_list
    # 测试结束恢复原始逻辑
    HttpReqCaptureProcessor.on_end = original_on_end

def test_business_logic(captured_http_requests):
    # 执行业务代码
    run_your_business_func()
    # 断言请求符合预期
    assert len(captured_http_requests) == 2
    assert captured_http_requests[0]["url"] == "https://api.example.com/expected_endpoint"

方案2:手动猴子补丁(轻量场景)

如果不想引入OpenTelemetry的依赖,你可以直接给标准库http.client的核心方法打补丁,覆盖所有基于标准库实现的HTTP客户端请求。

实现示例

import http.client
import logging
from functools import wraps

# 保存原始方法避免覆盖
original_request = http.client.HTTPConnection.request
original_getresponse = http.client.HTTPConnection.getresponse
# 存储请求上下文映射
_req_context = {}

@wraps(original_request)
def patched_request(self, method, url, body=None, headers={}, *, encode_chunked=False):
    req_id = id(self)
    # 记录请求基础信息
    _req_context[req_id] = {
        "method": method,
        "url": f"{'https' if self._http_vsn == 11 else 'http'}://{self.host}:{self.port}{url}",
        "body": body,
        "headers": headers
    }
    return original_request(self, method, url, body=body, headers=headers, encode_chunked=encode_chunked)

@wraps(original_getresponse)
def patched_getresponse(self):
    req_id = id(self)
    req_data = _req_context.pop(req_id, {})
    # 调用原始方法获取响应
    resp = original_getresponse(self)
    # 补全响应信息
    full_data = {
        **req_data,
        "status_code": resp.status,
        "response_headers": dict(resp.getheaders())
    }
    # 在此处添加自定义处理逻辑:日志、Kafka上报、DataFrame存储等
    logging.info(f"捕获HTTP请求: {full_data['method']} {full_data['url']} 状态码: {full_data['status_code']}")
    return resp

# 挂载补丁,注意要在所有业务代码导入前执行
http.client.HTTPConnection.request = patched_request
http.client.HTTPConnection.getresponse = patched_getresponse

注意事项

  • 若代码库使用aiohttp、httpx等异步HTTP客户端,需要单独给对应库的核心请求方法打补丁,原理和上面一致
  • 补丁必须在业务代码导入HTTP客户端之前执行,否则会出现部分请求漏捕获的情况

极端场景适配

如果代码库存在基于原生socket手动实现的HTTP请求,上述方案无法覆盖,你可以额外给socket.connect方法打补丁,过滤80、443端口的流量进行捕获。


内容的提问来源于stack exchange,提问作者Michael Wheeler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 07:57:02