You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自定义Locust Sagemaker User在定时压测达时长限制后仍持续运行问题

问题原因

1. 异常捕获逻辑吞掉了Locust的终止信号

你在predictEx方法中使用了裸except:语句,会捕获所有Python异常,包括Locust停止用户线程时抛出的StopUser专用异常。你捕获异常后仅上报了请求失败事件,没有重新抛出StopUser异常,导致Locust的用户终止逻辑被打断,用户线程会持续循环执行下一轮task,永远不会主动停止。

2. Sagemaker调用未配置超时限制

Sagemaker Predictor默认的同步调用超时时间较长(默认值为60秒),如果压测触发停止指令时刚好有请求已经发往Sagemaker且尚未返回,该调用会一直阻塞到超时才会结束,导致压测进程需要等待所有阻塞请求返回后才会完全终止,肉眼看起来就是任务仍在持续运行。

3. 未配置Locust停止超时参数

Locust默认收到停止指令后,会无限等待正在运行的task执行完成,没有最大等待时间限制。如果你的单个Sagemaker调用耗时很长,就会导致整体压测任务迟迟无法终止。


修复方案

  1. 修改异常捕获逻辑,不要吞掉Locust的终止异常:
from locust.exception import StopUser
# 在predictEx方法中修改异常捕获部分
def predictEx(self, data):
    start_time = time.time()
    start_perf_counter = time.perf_counter()
    name = 'predictEx'
    try:
        result = self.predict(data)
    except StopUser:
        # 收到停止信号直接抛出,不要吞掉
        raise
    except Exception:
        # 仅捕获业务请求相关异常,上报失败事件
        total_time = int((time.perf_counter() - start_perf_counter) * 1000)
        events.request_failure.fire(request_type="sagemaker", name=name, response_time=total_time, exception=sys.exc_info(), response_length=0)
    else:
        total_time = int((time.perf_counter() - start_perf_counter) * 1000)
        events.request_success.fire(request_type="sagemaker", name=name, response_time=total_time, response_length=sys.getsizeof(result))
  1. 初始化Sagemaker Predictor时配置调用超时,避免无限阻塞:
from sagemaker.predictor import PredictorConfig

self.client = SagemakerClient(
    sagemaker_session = Session(),
    endpoint_name = "sagemaker-test",
    serializer = JSONSerializer(),
    # 新增超时配置,连接和读取超时都设为10秒,可根据实际场景调整
    config=PredictorConfig(
        connect_timeout=10,
        read_timeout=10
    )
)
  1. 启动Locust时添加--stop-timeout参数,设置最大等待时间,超时后强制终止任务,比如设置最长等待15秒:
locust -f 你的脚本文件名.py --stop-timeout 15

内容的提问来源于stack exchange,提问作者cad86

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 03:27:03