You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure ML端点并行执行配置问题求助

解决Azure ML托管在线端点请求并行执行问题

问题分析

你配置了max_concurrent_requests_per_instance=4但仍出现串行处理,核心原因是底层Web服务器(默认是gunicorn)的并发参数未匹配Azure ML的设置。默认的minimal镜像默认并发配置较低,导致Azure ML的并发限制无法生效。

解决方案

无需修改评分脚本,只需调整部署环境的Web服务器参数,使其与max_concurrent_requests_per_instance设置匹配。具体步骤如下:

1. 调整环境变量配置

在创建Environment时,添加environment_variables设置gunicorn的工作进程和线程数,确保总并发数等于max_concurrent_requests_per_instance的值。

修改后的代码片段:

def create_online_deployment(self, ml_client, build_suffix, online_endpoint_name):
    env = Environment(
        conda_file='./env.yml',
        image='mcr.microsoft.com/azureml/minimal-ubuntu20.04-py38-cpu-inference:latest',
        # 添加gunicorn并发配置,总并发=workers*threads=2*2=4,匹配max_concurrent_requests_per_instance
        environment_variables={
            "GUNICORN_CMD_ARGS": "--workers=2 --threads=2 --timeout=120"
        }
    )
    deployment_name = f'deploy-{build_suffix}'
    lc_deployment = ManagedOnlineDeployment(
        name=deployment_name,
        environment=env,
        code_configuration=CodeConfiguration(
            code='./', scoring_script='ai_api/amp.py'
        ),
        request_settings=OnlineRequestSettings(
            request_timeout_ms=120000, 
            max_concurrent_requests_per_instance=4,
            max_queue_wait_ms=60000
        ),
        endpoint_name=online_endpoint_name,
        instance_type='Standard_E2s_v3',
        instance_count=1,
    )
    ml_client.online_deployments.begin_create_or_update(lc_deployment).result()
    self.logger.info(f'Created online deployment: {deployment_name}')

2. 参数说明

  • --workers=2:设置gunicorn工作进程数,建议等于实例的vCPU核心数(Standard_E2s_v3有2vCPU)。
  • --threads=2:每个工作进程的线程数,与workers相乘得到总并发数(2*2=4),和max_concurrent_requests_per_instance保持一致。
  • --timeout=120:设置gunicorn请求超时时间,需大于等于request_timeout_ms(120秒),避免请求被提前终止。

3. 验证方法

部署完成后,使用并发测试工具(如ab -n 10 -c 10 <你的端点URL>)发起10个并发请求,查看总耗时是否接近30秒,以此验证并行效果。

额外注意事项

  • 若实例规格变更(如使用更高vCPU的实例),需同步调整workers和threads的数值,确保总并发数与max_concurrent_requests_per_instance匹配。
  • 确保评分脚本的init()函数没有全局锁或阻塞操作,否则会影响并发执行(若存在此类问题,需最小化修改脚本消除阻塞,但你明确不想修改脚本,此点仅作提醒)。

内容的提问来源于stack exchange,提问作者yash pardeshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 17:52:46