新手调用Azure AI API遇大字典报错及请求限制问题求助
解决Azure AI API调用的三类问题:请求过大、限流、耗时过长
一、解决'400 Bad Request'(请求过大)问题
原因:将完整大字典直接塞入prompt,超过模型上下文窗口的token限制。你的字典包含4个key,每个key下2000条数据,转成文本后的token量远超常规模型(如gpt-3.5-turbo)的上限。
解决步骤:
- 拆分数据单元:按索引将数据拆分为单条或小批量的独立请求单元,只传入当前处理索引对应的4个key的value,而非整个字典。
- 控制单请求token量:根据所用模型的上下文限制(如gpt-3.5-turbo-16k支持16384 token),计算每次最多能传入的批量数据量,避免单次请求token超标。
示例数据拆分代码:
# 假设已从Excel读取得到my_dict data_rows = [] # 遍历每个索引,整理成单条处理单元 for idx in range(2000): single_row = { "key1": my_dict["key1"][idx], "key2": my_dict["key2"][idx], "key3": my_dict["key3"][idx], "key4": my_dict["key4"][idx] } data_rows.append(single_row)
修改后的单请求payload示例:
payload = { "messages": [ { "role": "system", "content": "你的系统提示词..." }, { "role": "user", "content": f"处理以下数据:{single_row}" } ], "temperature": 0.7 }
二、解决'429 Too Many Requests'(限流)问题
原因:Azure OpenAI对API调用有速率限制(每分钟请求数、每分钟token数),连续高频请求会触发限流机制。
解决步骤:
- 添加重试机制:捕获429错误后,根据响应头的
Retry-After值等待后重试,或使用指数退避策略。 - 控制请求速率:在请求之间添加固定或动态等待时间,确保不超过Azure资源的配额限制(可在Azure门户查看具体配额)。
示例带重试和限流的代码(使用tenacity库):
import requests import time from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type endpoint = 'https://resource-name.openai.azure.com/openai/deployments/deployment-id/chat/completions?api-version=2024-02-01' api_key = '你的API密钥' headers = {'Content-Type': 'application/json', 'api-key': api_key} @retry( stop=stop_after_attempt(5), # 最多重试5次 wait=wait_exponential(multiplier=1, min=2, max=10), # 指数退避:2s→4s→8s…最多10s retry=retry_if_exception_type(requests.exceptions.HTTPError) ) def call_azure_api(payload): response = requests.post(endpoint, headers=headers, json=payload) response.raise_for_status() # 抛出HTTP错误触发重试 return response.json() # 遍历处理数据,每次请求后等待1秒(可根据配额调整) results = [] for row in data_rows[:10]: # 先测试前10条 payload = { "messages": [ {"role": "system", "content": "你的系统提示词..."}, {"role": "user", "content": f"处理以下数据:{row}"} ], "temperature": 0.7 } try: result = call_azure_api(payload) results.append(result) time.sleep(1) except Exception as e: print(f"处理数据失败:{e}") results.append(None)
三、解决小样本处理耗时过长问题
原因:同步请求会阻塞等待响应,单线程处理效率低下。
解决步骤:
- 使用异步请求:用
aiohttp替代requests,实现多并发请求,同时控制并发数避免触发限流。 - 批量请求优化:根据模型支持情况,调整批量处理的大小,减少请求总次数。
示例异步请求代码:
import aiohttp import asyncio endpoint = 'https://resource-name.openai.azure.com/openai/deployments/deployment-id/chat/completions?api-version=2024-02-01' api_key = '你的API密钥' headers = {'Content-Type': 'application/json', 'api-key': api_key} async def async_call_api(session, payload): async with session.post(endpoint, headers=headers, json=payload) as response: if response.status == 429: retry_after = int(response.headers.get('Retry-After', 2)) await asyncio.sleep(retry_after) return await async_call_api(session, payload) response.raise_for_status() return await response.json() async def process_batch(data_batch): async with aiohttp.ClientSession() as session: tasks = [] for row in data_batch: payload = { "messages": [ {"role": "system", "content": "你的系统提示词..."}, {"role": "user", "content": f"处理以下数据:{row}"} ], "temperature": 0.7 } tasks.append(async_call_api(session, payload)) # 控制并发数,一次处理5个请求 results = await asyncio.gather(*tasks, return_exceptions=True) return results # 批量处理,每次处理5条数据 batch_size = 5 all_results = [] for i in range(0, len(data_rows), batch_size): batch = data_rows[i:i+batch_size] batch_results = await process_batch(batch) all_results.extend(batch_results) await asyncio.sleep(2) # 批量间添加等待时间 # 运行异步代码(本地环境直接执行,Jupyter需调整事件循环) asyncio.run(process_batch(data_rows[:10]))
额外优化建议
- 简化prompt:去掉冗余描述,用简洁指令减少token消耗,同时提升处理速度。
- 查看Azure配额:在Azure门户的OpenAI资源页面,确认请求速率和token配额,根据配额调整请求频率和批量大小。
- 使用批量API:若场景适合,可使用Azure OpenAI的批量处理API,打包所有请求上传后后台处理,避免频繁请求。
内容的提问来源于stack exchange,提问作者Iluvatar Bombadil
相关产品推荐
相关产品推荐

