You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure API Management如何实现三后端的负载均衡与故障转移?

解决方案:Azure API Management 三后端故障转移+轮询负载均衡

首先明确:Azure APIM的<retry>策略中没有内置的context.RetryCount变量,你需要通过自定义上下文变量来追踪重试/后端切换次数,以下是两种可行方案:

方案一:基础故障转移(按顺序循环重试)

该方案实现核心需求:backend1失败→backend2→backend3→backend1循环,直到重试次数耗尽或请求成功。

<policies>
    <inbound>
        <!-- 初始化后端索引:0=backend1,1=backend2,2=backend3 -->
        <set-variable name="currentBackendIndex" value="@(0)" />
    </inbound>
    <backend>
        <!-- 首次请求绑定默认后端(backend1) -->
        <set-backend-service backend-id="@(new string[]{"backend1", "backend2", "backend3"}[(int)context.Variables["currentBackendIndex"]])" />
        <forward-request />
        
        <!-- 重试触发条件:后端返回5xx服务器错误 -->
        <retry condition="@(context.Response != null && (int)context.Response.StatusCode >= 500)" count="9" interval="10" first-fast-retry="true">
            <!-- 递增索引并取模3,实现后端循环切换 -->
            <set-variable name="currentBackendIndex" value="@(((int)context.Variables["currentBackendIndex"] + 1) % 3)" />
            
            <!-- 切换到下一个后端 -->
            <set-backend-service backend-id="@(new string[]{"backend1", "backend2", "backend3"}[(int)context.Variables["currentBackendIndex"]])" />
            
            <forward-request />
        </retry>
    </backend>
    <outbound>
    </outbound>
    <on-error>
    </on-error>
</policies>

关键逻辑说明

  • 首次请求直接使用backend1,失败后触发重试流程
  • 每次重试时自动切换到下一个后端,索引取模3实现循环
  • count="9"表示最多重试9次,加上首次请求共尝试10次,可根据业务需求调整
  • first-fast-retry="true"表示第一次重试不等待interval设置的10秒,后续重试间隔10秒

方案二:故障转移+轮询负载均衡

如果需要在正常请求时实现轮询负载均衡,失败时自动故障转移,可以结合APIM内置缓存维护全局轮询索引:

<policies>
    <inbound>
        <!-- 从全局缓存获取当前轮询索引,默认初始为0 -->
        <cache-lookup-value key="globalBackendIndex" variable-name="globalIndex" default-value="0" />
        <!-- 更新全局索引,取模3实现轮询,缓存有效期1秒避免并发冲突 -->
        <cache-store-value key="globalBackendIndex" value="@(((int)context.Variables["globalIndex"] + 1) % 3)" duration="1" />
        <!-- 将当前请求的后端索引存入上下文变量 -->
        <set-variable name="currentBackendIndex" value="@((int)context.Variables["globalIndex"])" />
    </inbound>
    <backend>
        <!-- 绑定轮询分配的初始后端 -->
        <set-backend-service backend-id="@(new string[]{"backend1", "backend2", "backend3"}[(int)context.Variables["currentBackendIndex"]])" />
        <forward-request />
        
        <!-- 重试触发条件:覆盖5xx错误及请求超时 -->
        <retry condition="@(context.Response != null && ((int)context.Response.StatusCode >= 500 || (int)context.Response.StatusCode == 408))" count="9" interval="10" first-fast-retry="true">
            <!-- 失败时切换到下一个后端 -->
            <set-variable name="currentBackendIndex" value="@(((int)context.Variables["currentBackendIndex"] + 1) % 3)" />
            <set-backend-service backend-id="@(new string[]{"backend1", "backend2", "backend3"}[(int)context.Variables["currentBackendIndex"]])" />
            
            <forward-request />
        </retry>
    </backend>
    <outbound>
    </outbound>
    <on-error>
    </on-error>
</policies>

关键逻辑说明

  • 正常请求通过全局缓存实现轮询分配后端,每个请求按顺序轮流使用backend1/backend2/backend3
  • 当某个后端失败时,自动切换到下一个后端重试,循环往复
  • 重试条件扩展了408请求超时,可根据实际需要调整错误码范围

内容的提问来源于stack exchange,提问作者Rafferty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 23:30:38