You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Locust追踪API响应的token_spend指标并聚合分析延迟

用Locust实现全局Token消耗统计与延迟对比分析

1. 全局Token消耗聚合

Locust没有内置全局统计能力,需自行实现线程安全的全局计数器:

  • 定义全局变量存储总Token消耗,用threading.Lock保证并发场景下的数据安全
  • 通过events.request_success事件钩子,在每个请求成功后从响应中提取token_spend值,累加至全局计数器

2. 记录每分钟Token消耗速率

用定时器定期计算并记录速率:

  • 维护变量记录上一次统计的Token总量
  • 每隔60秒触发一次统计,用当前总Token减去上次值得到每分钟消耗,将结果写入日志或文件
  • 在测试启动时初始化定时器,测试结束时关闭日志文件

3. 不同Token消耗区间的延迟对比

将每个请求的Token消耗与延迟关联,按区间分组统计:

  • 预设Token区间(比如0-10、10-20、20+),用字典存储每个区间的延迟数据列表
  • 在events.request_success钩子中,根据当前请求的token_spend值判断所属区间,将延迟值加入对应列表
  • 测试结束时,对每个区间的延迟数据计算平均值、中位数等指标,完成对比分析

完整代码示例

from locust import HttpUser, task, events
import threading
import time

# 全局Token统计相关
total_token_spend = 0
token_lock = threading.Lock()
last_token_count = 0
rate_log_file = open("token_rate.log", "w")

# Token区间延迟统计
token_intervals = {
    "0-10": [],
    "10-20": [],
    "20+": []
}
interval_lock = threading.Lock()

def calculate_token_rate():
    global total_token_spend, last_token_count
    with token_lock:
        current = total_token_spend
        rate = current - last_token_count
        last_token_count = current
    rate_log_file.write(f"{time.strftime('%Y-%m-%d %H:%M:%S')}, 每分钟Token消耗: {rate}\n")
    rate_log_file.flush()
    # 递归调用,持续统计
    threading.Timer(60, calculate_token_rate).start()

@events.test_start.add_listener
def on_test_start(environment, **kwargs):
    # 启动每分钟统计定时器
    calculate_token_rate()

@events.test_stop.add_listener
def on_test_stop(environment, **kwargs):
    rate_log_file.close()
    # 输出各Token区间的延迟统计
    print("\n=== Token消耗区间延迟对比 ===")
    for interval, latencies in token_intervals.items():
        if latencies:
            avg_latency = sum(latencies) / len(latencies)
            median_latency = sorted(latencies)[len(latencies)//2]
            print(f"区间 {interval}: 平均延迟 {avg_latency:.2f}ms, 中位数延迟 {median_latency:.2f}ms")

@events.request_success.add_listener
def on_request_success(request_type, name, response_time, response_length, response, **kwargs):
    global total_token_spend
    # 提取token_spend字段
    try:
        token_spend = response.json().get("token_spend", 0)
    except:
        token_spend = 0

    # 更新全局Token总量
    with token_lock:
        total_token_spend += token_spend

    # 更新对应Token区间的延迟数据
    with interval_lock:
        if token_spend <= 10:
            token_intervals["0-10"].append(response_time)
        elif token_spend <= 20:
            token_intervals["10-20"].append(response_time)
        else:
            token_intervals["20+"].append(response_time)

class APITestUser(HttpUser):
    @task
    def test_api(self):
        self.client.get("/your-api-endpoint")
        # 响应处理由钩子自动完成,无需额外操作

注意事项

  • 分布式测试时,本地全局计数器会失效,需改用Redis等共享存储来汇总各Worker的Token数据
  • 若API响应格式不同,需调整response.json().get("token_spend")的提取逻辑
  • 测试结束后,可将token_rate.log和区间延迟统计结果导出,做进一步可视化分析

内容的提问来源于stack exchange,提问作者Mouse Rat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 13:32:57