Azure OpenAI+APIM负载均衡配置后,Python代码如何适配?
问题
已完成Azure API Management(APIM)与Azure OpenAI的负载均衡策略配置,搭建了可生成文本摘要的基础Web应用。此前代码直接调用Azure OpenAI端点并使用密钥,现需修改Python代码以适配APIM负载均衡端点。
附APIM策略代码:
<!-- IMPORTANT: - Policy elements can appear only within the <inbound>, <outbound>, <backend> section elements. - To apply a policy to the incoming request (before it is forwarded to the backend service), place a corresponding policy element within the <inbound> section element. - To apply a policy to the outgoing response (before it is sent back to the caller), place a corresponding policy element within the <outbound> section element. - To add a policy, place the cursor at the desired insertion point and select a policy from the sidebar. - To remove a policy, delete the corresponding policy statement from the policy document. - Position the <base> element within a section element to inherit all policies from the corresponding section element in the enclosing scope. - Remove the <base> element to prevent inheriting policies from the corresponding section element in the enclosing scope. - Policies are applied in the order of their appearance, from the top down. - Comments within policy elements are not supported and may disappear. Place your comments between policy elements or at a higher level scope. --> <policies> <inbound> <base /> <check-header name="X-Azure-FDID" failed-check-httpcode="401" failed-check-error-message="Not authorized" ignore-case="false"> <value>aaabbbbccc</value> </check-header> <ip-filter action="allow"> <address-range from="add1" to="add2" /> </ip-filter> <set-variable name="backendUrlA" value="url1" /> <!-- [A] Subscription 1: location East --> <set-variable name="backendUrlB" value="url2" /> <!-- [B] Subscription 1: location East --> <set-variable name="backendA-apiKey" value="{{backendA-apiKey}}" /> <set-variable name="backendB-apiKey" value="{{backendB-apiKey}}" /> <!-- Load balancing logic --> <choose> <when condition="@((int)context.Request.Url.Path.IndexOf("/gpt-35-turbo") != -1)"> <!-- Pool 1: GPT3.5 Turbo and GPT3.5 Turbo 16K --> <!-- Load balancing logic for gpt-35-turbo models --> <cache-lookup-value key="pool1Index" default-value="@((int)0)" variable-name="pool1Index" /> <choose> <when condition="@( (int)context.Variables["pool1Index"] == 0 )"> <set-variable name="selectedBackend" value="@((string)context.Variables["backendUrlA"])" /> <set-variable name="selectedBackendKey" value="@((string)context.Variables["backendA-apiKey"])" /> </when> <otherwise> <set-variable name="selectedBackend" value="@((string)context.Variables["backendUrlB"])" /> <set-variable name="selectedBackendKey" value="@((string)context.Variables["backendB-apiKey"])" /> </otherwise> </choose> <!-- Increment the pool1Index and reset to 0 when it reaches the end --> <set-variable name="pool1Index" value="@((int)context.Variables["pool1Index"] == 1 ? 0 : (int)context.Variables["pool1Index"] + 1)" /> <cache-store-value key="pool1Index" value="@((int)context.Variables["pool1Index"])" duration="1440" /> </when> <when condition="@((int)context.Request.Url.Path.IndexOf("/text-embedding-ada-002") != -1)"> <!-- Pool 2: Embedding ADA-002 --> <!-- Load balancing logic for text-embedding-ada-002 model --> <cache-lookup-value key="pool2Index" default-value="@((int)0)" variable-name="pool2Index" /> <choose> <when condition="@( (int)context.Variables["pool2Index"] == 0 )"> <set-variable name="selectedBackend" value="@((string)context.Variables["backendUrlA"])" /> <set-variable name="selectedBackendKey" value="@((string)context.Variables["backendA-apiKey"])" /> </when> <otherwise> <set-variable name="selectedBackend" value="@((string)context.Variables["backendUrlB"])" /> <set-variable name="selectedBackendKey" value="@((string)context.Variables["backendB-apiKey"])" /> </otherwise> </choose> <!-- Increment the pool2Index and reset to 0 when it reaches the end --> <set-variable name="pool2Index" value="@((int)context.Variables["pool2Index"] == 1 ? 0 : (int)context.Variables["pool2Index"] + 1)" /> <cache-store-value key="pool2Index" value="@((int)context.Variables["pool2Index"])" duration="1440" /> </when> <otherwise> <!-- Direct to Pool 3: General --> <set-variable name="selectedBackend" value="@((string)context.Variables["backendUrlA"])" /> <set-variable name="selectedBackendKey" value="@((string)context.Variables["backendA-apiKey"])" /> </otherwise> </choose> <set-backend-service base-url="@((string)context.Variables["selectedBackend"])" /> <set-header name="api-key" exists-action="override"> <value>@((string)context.Variables["selectedBackendKey"])</value> </set-header> </inbound> <backend> <base /> </backend> <outbound> <base /> <!-- Add the selected backend URL to the response headers --> <set-header name="X-Selected-Backend" exists-action="override"> <value>@((string)context.Variables["selectedBackend"])</value> </set-header> </outbound> <on-error> <base /> </on-error> </policies>
解决方案
核心修改点
根据你的APIM策略,代码需要做以下关键调整:
- 替换请求端点为APIM的服务端点(而非直接调用Azure OpenAI端点)
- 添加APIM强制验证的
X-Azure-FDID请求头(值为策略中指定的aaabbbbccc) - 移除原有的Azure OpenAI密钥配置(APIM会自动注入后端实例的密钥)
- 保持原有请求体结构不变(APIM会转发请求到对应后端)
代码示例1:使用OpenAI Python SDK
from openai import AzureOpenAI # 初始化客户端时指定APIM端点,无需设置api_key client = AzureOpenAI( azure_endpoint="https://你的APIM实例名称.azure-api.net/openai", # 替换为你的APIM端点 api_version="2024-02-15-preview", # 保持与原代码一致的API版本 # 注意:不需要设置api_key,APIM会自动处理后端密钥 ) # 构造请求时添加X-Azure-FDID头 response = client.chat.completions.create( model="gpt-35-turbo", messages=[ {"role": "system", "content": "请生成文本摘要"}, {"role": "user", "content": "你的输入文本内容"} ], extra_headers={ "X-Azure-FDID": "aaabbbbccc" # 必须与APIM策略中的值一致 } ) # 可选:查看响应头中的X-Selected-Backend,验证负载均衡是否生效 print("当前选中的后端实例:", response.headers.get("X-Selected-Backend")) print("摘要结果:", response.choices[0].message.content)
代码示例2:使用requests库直接调用
import requests APIM_ENDPOINT = "https://你的APIM实例名称.azure-api.net/openai/deployments/gpt-35-turbo/chat/completions" # 替换为你的APIM完整端点 API_VERSION = "2024-02-15-preview" headers = { "Content-Type": "application/json", "X-Azure-FDID": "aaabbbbccc", # 必须添加的验证头 # 注意:不需要添加api-key头,APIM会自动注入 } payload = { "messages": [ {"role": "system", "content": "请生成文本摘要"}, {"role": "user", "content": "你的输入文本内容"} ], "temperature": 0.7 } response = requests.post( f"{APIM_ENDPOINT}?api-version={API_VERSION}", headers=headers, json=payload ) # 检查响应状态 if response.status_code == 200: result = response.json() print("当前选中的后端实例:", response.headers.get("X-Selected-Backend")) print("摘要结果:", result["choices"][0]["message"]["content"]) else: print(f"请求失败: {response.status_code} - {response.text}")
关键说明
- IP白名单验证:确保你的Web应用服务器IP在APIM策略的
<ip-filter>允许范围内,否则会被拒绝访问。 - 负载均衡验证:通过响应头
X-Selected-Backend可以查看当前请求被路由到哪个Azure OpenAI实例,确认负载均衡逻辑是否正常工作。 - 模型路径匹配:APIM策略通过请求路径中的模型名称(如
/gpt-35-turbo)分配后端池,确保请求URL中的模型名称与策略中的匹配。
内容的提问来源于stack exchange,提问作者FarzadAvari
相关产品推荐
相关产品推荐

