You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure OpenAI部署APIM后端切换逻辑问题排查与优化请求

优化Azure APIM动态切换Azure OpenAI部署以避免429错误的方案

需求背景

  • 主部署:PTU实例,配置每分钟10000令牌配额
  • 备用部署:Pay-per-go实例,触发切换的条件为:
    • PTU剩余令牌≤600
    • PTU返回429状态码

当前问题

并行请求场景下,APIM的令牌剩余量仅在请求成功执行后更新,导致多请求同时消耗PTU令牌时,还未触发剩余量≤600的切换条件,令牌就已耗尽并触发429错误,切换逻辑滞后。

优化方案

1. 提前预估令牌消耗,前置切换判断

启用APIM的令牌预估功能,在入站阶段提前计算当前请求的令牌消耗,预判剩余令牌扣除本次消耗后的余量,提前触发切换逻辑,避免并行请求击穿阈值。

2. 全局标记PTU超限状态

当PTU返回429时,通过APIM缓存设置全局超限标记,后续请求直接跳过PTU检查,切换到备用部署,直到缓存过期(与响应的Retry-After时间同步)。

3. 调整阈值判断逻辑

将原有的"剩余令牌≤600"判断,改为"剩余令牌-预估消耗≤600",确保当前请求的令牌消耗不会导致剩余量跌破安全线。

修改后的策略代码

入站策略

<inbound>
    <base />
    <choose>
        <when condition="@(context.Request.Url.Path.Contains('/chat/completions'))">
            <!-- 检查全局缓存中PTU是否已超限 -->
            <cache-lookup-value key="PTU_LIMIT_EXCEEDED" variable-name="ptuLimitExceeded" />
            <choose>
                <when condition="@(context.Variables.GetValueOrDefault<bool>("ptuLimitExceeded"))">
                    <set-variable name="backendUrl" value="https://pay-per-go-backend-url.com/" />
                    <set-variable name="backendApiKey" value="PAY_PER_GO_API_KEY" />
                </when>
                <otherwise>
                    <!-- 启用令牌预估,获取本次请求的预估消耗 -->
                    <azure-openai-token-limit 
                        counter-key="@(context.Request.Headers.GetValueOrDefault('Ocp-Apim-Subscription-Key'))" 
                        tokens-per-minute="10000" 
                        estimate-prompt-tokens="true" 
                        remaining-tokens-variable-name="remainingTokens"
                        consumed-tokens-variable-name="consumedTokens" />
                    <!-- 计算扣除本次消耗后的可用令牌 -->
                    <set-variable name="availableTokens" value="@(context.Variables.GetValueOrDefault<int>("remainingTokens") - context.Variables.GetValueOrDefault<int>("consumedTokens"))" />
                    <choose>
                        <!-- 可用令牌不足时切换到备用部署 -->
                        <when condition="@(context.Variables.ContainsKey("availableTokens") && context.Variables.GetValueOrDefault<int>("availableTokens") <= 600)">
                            <set-variable name="backendUrl" value="https://pay-per-go-backend-url.com/" />
                            <set-variable name="backendApiKey" value="PAY_PER_GO_API_KEY" />
                        </when>
                        <otherwise>
                            <set-variable name="backendUrl" value="https://ptu-url.com/" />
                            <set-variable name="backendApiKey" value="PTU_API_KEY" />
                        </otherwise>
                    </choose>
                </otherwise>
            </choose>
        </when>
    </choose>
</inbound>

后端策略

<backend>
    <choose>
        <when condition="@(context.Response.StatusCode == 429 && context.Variables.GetValueOrDefault<string>("backendUrl") == "https://ptu-url.com/")">
            <!-- 解析Retry-After时间,默认60秒 -->
            <set-variable name="retryAfter" value="@(int.TryParse(context.Response.Headers.GetValueOrDefault("Retry-After"), out int retry) ? retry : 60)" />
            <!-- 缓存PTU超限标记,过期时间与Retry-After同步 -->
            <cache-store-value key="PTU_LIMIT_EXCEEDED" value="true" duration="@(context.Variables.GetValueOrDefault<int>("retryAfter"))" />
            <!-- 切换到备用部署并重试请求 -->
            <set-variable name="backendUrl" value="https://pay-per-go-backend-url.com/" />
            <set-variable name="backendApiKey" value="PAY_PER_GO_API_KEY" />
            <set-backend-service base-url="@(context.Variables.GetValueOrDefault<string>("backendUrl"))" />
            <forward-request buffer-request-body="true" />
        </when>
        <otherwise>
            <set-backend-service base-url="@(context.Variables.GetValueOrDefault<string>("backendUrl"))" />
            <forward-request buffer-request-body="true" />
        </otherwise>
    </choose>
</backend>

优化说明

  • 令牌预估:启用estimate-prompt-tokens="true"后,APIM会提前解析请求体计算令牌消耗,在入站阶段即可预判剩余令牌是否足够支撑当前请求,避免并行请求同时消耗导致超限。
  • 全局缓存标记:通过缓存存储PTU超限状态,确保后续请求直接走备用部署,直到配额重置,避免重复触发429错误。
  • 缓冲阈值:基于"剩余令牌-预估消耗"的判断逻辑,预留足够缓冲空间,防止单请求消耗直接击穿阈值。

内容的提问来源于stack exchange,提问作者silenthunter25

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 01:23:16