Azure ML端点并行执行配置问题求助
解决Azure ML托管在线端点请求并行执行问题
问题分析
你配置了max_concurrent_requests_per_instance=4但仍出现串行处理,核心原因是底层Web服务器(默认是gunicorn)的并发参数未匹配Azure ML的设置。默认的minimal镜像默认并发配置较低,导致Azure ML的并发限制无法生效。
解决方案
无需修改评分脚本,只需调整部署环境的Web服务器参数,使其与max_concurrent_requests_per_instance设置匹配。具体步骤如下:
1. 调整环境变量配置
在创建Environment时,添加environment_variables设置gunicorn的工作进程和线程数,确保总并发数等于max_concurrent_requests_per_instance的值。
修改后的代码片段:
def create_online_deployment(self, ml_client, build_suffix, online_endpoint_name): env = Environment( conda_file='./env.yml', image='mcr.microsoft.com/azureml/minimal-ubuntu20.04-py38-cpu-inference:latest', # 添加gunicorn并发配置,总并发=workers*threads=2*2=4,匹配max_concurrent_requests_per_instance environment_variables={ "GUNICORN_CMD_ARGS": "--workers=2 --threads=2 --timeout=120" } ) deployment_name = f'deploy-{build_suffix}' lc_deployment = ManagedOnlineDeployment( name=deployment_name, environment=env, code_configuration=CodeConfiguration( code='./', scoring_script='ai_api/amp.py' ), request_settings=OnlineRequestSettings( request_timeout_ms=120000, max_concurrent_requests_per_instance=4, max_queue_wait_ms=60000 ), endpoint_name=online_endpoint_name, instance_type='Standard_E2s_v3', instance_count=1, ) ml_client.online_deployments.begin_create_or_update(lc_deployment).result() self.logger.info(f'Created online deployment: {deployment_name}')
2. 参数说明
--workers=2:设置gunicorn工作进程数,建议等于实例的vCPU核心数(Standard_E2s_v3有2vCPU)。--threads=2:每个工作进程的线程数,与workers相乘得到总并发数(2*2=4),和max_concurrent_requests_per_instance保持一致。--timeout=120:设置gunicorn请求超时时间,需大于等于request_timeout_ms(120秒),避免请求被提前终止。
3. 验证方法
部署完成后,使用并发测试工具(如ab -n 10 -c 10 <你的端点URL>)发起10个并发请求,查看总耗时是否接近30秒,以此验证并行效果。
额外注意事项
- 若实例规格变更(如使用更高vCPU的实例),需同步调整
workers和threads的数值,确保总并发数与max_concurrent_requests_per_instance匹配。 - 确保评分脚本的
init()函数没有全局锁或阻塞操作,否则会影响并发执行(若存在此类问题,需最小化修改脚本消除阻塞,但你明确不想修改脚本,此点仅作提醒)。
内容的提问来源于stack exchange,提问作者yash pardeshi
相关产品推荐
相关产品推荐

