You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新手调用Azure AI API遇大字典报错及请求限制问题求助

解决Azure AI API调用的三类问题:请求过大、限流、耗时过长

一、解决'400 Bad Request'(请求过大)问题

原因:将完整大字典直接塞入prompt,超过模型上下文窗口的token限制。你的字典包含4个key,每个key下2000条数据,转成文本后的token量远超常规模型(如gpt-3.5-turbo)的上限。

解决步骤:

  • 拆分数据单元:按索引将数据拆分为单条或小批量的独立请求单元,只传入当前处理索引对应的4个key的value,而非整个字典。
  • 控制单请求token量:根据所用模型的上下文限制(如gpt-3.5-turbo-16k支持16384 token),计算每次最多能传入的批量数据量,避免单次请求token超标。

示例数据拆分代码:

# 假设已从Excel读取得到my_dict
data_rows = []
# 遍历每个索引,整理成单条处理单元
for idx in range(2000):
    single_row = {
        "key1": my_dict["key1"][idx],
        "key2": my_dict["key2"][idx],
        "key3": my_dict["key3"][idx],
        "key4": my_dict["key4"][idx]
    }
    data_rows.append(single_row)

修改后的单请求payload示例:

payload = {
    "messages": [
        {
            "role": "system",
            "content": "你的系统提示词..."
        },
        {
            "role": "user",
            "content": f"处理以下数据:{single_row}"
        }
    ],
    "temperature": 0.7
}

二、解决'429 Too Many Requests'(限流)问题

原因:Azure OpenAI对API调用有速率限制(每分钟请求数、每分钟token数),连续高频请求会触发限流机制。

解决步骤:

  • 添加重试机制:捕获429错误后,根据响应头的Retry-After值等待后重试,或使用指数退避策略。
  • 控制请求速率:在请求之间添加固定或动态等待时间,确保不超过Azure资源的配额限制(可在Azure门户查看具体配额)。

示例带重试和限流的代码(使用tenacity库):

import requests
import time
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

endpoint = 'https://resource-name.openai.azure.com/openai/deployments/deployment-id/chat/completions?api-version=2024-02-01'
api_key = '你的API密钥'
headers = {'Content-Type': 'application/json', 'api-key': api_key}

@retry(
    stop=stop_after_attempt(5),  # 最多重试5次
    wait=wait_exponential(multiplier=1, min=2, max=10),  # 指数退避:2s→4s→8s…最多10s
    retry=retry_if_exception_type(requests.exceptions.HTTPError)
)
def call_azure_api(payload):
    response = requests.post(endpoint, headers=headers, json=payload)
    response.raise_for_status()  # 抛出HTTP错误触发重试
    return response.json()

# 遍历处理数据,每次请求后等待1秒(可根据配额调整)
results = []
for row in data_rows[:10]:  # 先测试前10条
    payload = {
        "messages": [
            {"role": "system", "content": "你的系统提示词..."},
            {"role": "user", "content": f"处理以下数据:{row}"}
        ],
        "temperature": 0.7
    }
    try:
        result = call_azure_api(payload)
        results.append(result)
        time.sleep(1)
    except Exception as e:
        print(f"处理数据失败:{e}")
        results.append(None)

三、解决小样本处理耗时过长问题

原因:同步请求会阻塞等待响应,单线程处理效率低下。

解决步骤:

  • 使用异步请求:用aiohttp替代requests,实现多并发请求,同时控制并发数避免触发限流。
  • 批量请求优化:根据模型支持情况,调整批量处理的大小,减少请求总次数。

示例异步请求代码:

import aiohttp
import asyncio

endpoint = 'https://resource-name.openai.azure.com/openai/deployments/deployment-id/chat/completions?api-version=2024-02-01'
api_key = '你的API密钥'
headers = {'Content-Type': 'application/json', 'api-key': api_key}

async def async_call_api(session, payload):
    async with session.post(endpoint, headers=headers, json=payload) as response:
        if response.status == 429:
            retry_after = int(response.headers.get('Retry-After', 2))
            await asyncio.sleep(retry_after)
            return await async_call_api(session, payload)
        response.raise_for_status()
        return await response.json()

async def process_batch(data_batch):
    async with aiohttp.ClientSession() as session:
        tasks = []
        for row in data_batch:
            payload = {
                "messages": [
                    {"role": "system", "content": "你的系统提示词..."},
                    {"role": "user", "content": f"处理以下数据:{row}"}
                ],
                "temperature": 0.7
            }
            tasks.append(async_call_api(session, payload))
        # 控制并发数,一次处理5个请求
        results = await asyncio.gather(*tasks, return_exceptions=True)
        return results

# 批量处理,每次处理5条数据
batch_size = 5
all_results = []
for i in range(0, len(data_rows), batch_size):
    batch = data_rows[i:i+batch_size]
    batch_results = await process_batch(batch)
    all_results.extend(batch_results)
    await asyncio.sleep(2)  # 批量间添加等待时间

# 运行异步代码(本地环境直接执行,Jupyter需调整事件循环)
asyncio.run(process_batch(data_rows[:10]))

额外优化建议

  • 简化prompt:去掉冗余描述,用简洁指令减少token消耗,同时提升处理速度。
  • 查看Azure配额:在Azure门户的OpenAI资源页面,确认请求速率和token配额,根据配额调整请求频率和批量大小。
  • 使用批量API:若场景适合,可使用Azure OpenAI的批量处理API,打包所有请求上传后后台处理,避免频繁请求。

内容的提问来源于stack exchange,提问作者Iluvatar Bombadil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 21:35:04