自定义Locust Sagemaker User在定时压测达时长限制后仍持续运行问题
问题原因
1. 异常捕获逻辑吞掉了Locust的终止信号
你在predictEx方法中使用了裸except:语句,会捕获所有Python异常,包括Locust停止用户线程时抛出的StopUser专用异常。你捕获异常后仅上报了请求失败事件,没有重新抛出StopUser异常,导致Locust的用户终止逻辑被打断,用户线程会持续循环执行下一轮task,永远不会主动停止。
2. Sagemaker调用未配置超时限制
Sagemaker Predictor默认的同步调用超时时间较长(默认值为60秒),如果压测触发停止指令时刚好有请求已经发往Sagemaker且尚未返回,该调用会一直阻塞到超时才会结束,导致压测进程需要等待所有阻塞请求返回后才会完全终止,肉眼看起来就是任务仍在持续运行。
3. 未配置Locust停止超时参数
Locust默认收到停止指令后,会无限等待正在运行的task执行完成,没有最大等待时间限制。如果你的单个Sagemaker调用耗时很长,就会导致整体压测任务迟迟无法终止。
修复方案
- 修改异常捕获逻辑,不要吞掉Locust的终止异常:
from locust.exception import StopUser # 在predictEx方法中修改异常捕获部分 def predictEx(self, data): start_time = time.time() start_perf_counter = time.perf_counter() name = 'predictEx' try: result = self.predict(data) except StopUser: # 收到停止信号直接抛出,不要吞掉 raise except Exception: # 仅捕获业务请求相关异常,上报失败事件 total_time = int((time.perf_counter() - start_perf_counter) * 1000) events.request_failure.fire(request_type="sagemaker", name=name, response_time=total_time, exception=sys.exc_info(), response_length=0) else: total_time = int((time.perf_counter() - start_perf_counter) * 1000) events.request_success.fire(request_type="sagemaker", name=name, response_time=total_time, response_length=sys.getsizeof(result))
- 初始化Sagemaker Predictor时配置调用超时,避免无限阻塞:
from sagemaker.predictor import PredictorConfig self.client = SagemakerClient( sagemaker_session = Session(), endpoint_name = "sagemaker-test", serializer = JSONSerializer(), # 新增超时配置,连接和读取超时都设为10秒,可根据实际场景调整 config=PredictorConfig( connect_timeout=10, read_timeout=10 ) )
- 启动Locust时添加
--stop-timeout参数,设置最大等待时间,超时后强制终止任务,比如设置最长等待15秒:
locust -f 你的脚本文件名.py --stop-timeout 15
内容的提问来源于stack exchange,提问作者cad86
相关产品推荐
相关产品推荐

