使用OpenAI API生成Embedding时遭遇服务器过载错误求助
解决OpenAI Embedding请求时的服务器过载错误
问题描述
批量生成HPO描述的Embedding时,收到服务器过载错误:
The server is currently overloaded with other requests. Sorry about that! You can retry your request, or contact us through our help center at help.openai.com if the error persists.
你的代码如下:
# 从JSON文件加载HPO字典 with open("hpos.json") as f: hpos_dict = json.load(f) # 创建存储embeddings的字典 hpo_embeddings = {} i = 0 hposNumber = len(hpos_dict) # 生成embedding并存储到字典中 for hpo_id, hpo_descs in hpos_dict.items(): embedding_list = [] for hpo_desc in hpo_descs: response = openai.Embedding.create( input=hpo_desc, model="text-embedding-ada-002" ) embedding_list.append(response["data"][0]["embedding"]) hpo_embeddings[hpo_id] = embedding_list i = i + 1 print( str(i) + "/" + str(hposNumber) ) # 将embedding字典保存到JSON文件 with open("hpo_embeddings.json", "w") as f: json.dump(hpo_embeddings, f)
解决方案
这个错误是因为OpenAI服务器负载过高,加上代码采用同步串行请求,短时间内发送大量请求触发了服务器保护机制。可以通过以下几种方式解决:
1. 添加指数退避重试机制
遇到过载错误时自动重试,且每次重试间隔逐渐延长,避免持续给服务器施压。示例代码:
import time import json import openai # 从JSON文件加载HPO字典 with open("hpos.json") as f: hpos_dict = json.load(f) # 创建存储embeddings的字典 hpo_embeddings = {} i = 0 hposNumber = len(hpos_dict) # 生成embedding并存储到字典中 for hpo_id, hpo_descs in hpos_dict.items(): embedding_list = [] for hpo_desc in hpo_descs: retry_count = 0 max_retries = 5 while retry_count < max_retries: try: response = openai.Embedding.create( input=hpo_desc, model="text-embedding-ada-002" ) embedding_list.append(response["data"][0]["embedding"]) break except openai.error.APIError as e: if "overloaded" in str(e).lower(): retry_count += 1 wait_time = 2 ** retry_count # 指数退避:2,4,8,16,32秒 print(f"服务器过载,{wait_time}秒后重试...") time.sleep(wait_time) else: raise e except Exception as e: print(f"未知错误:{e}") raise e hpo_embeddings[hpo_id] = embedding_list i = i + 1 print(f"{i}/{hposNumber}") # 将embedding字典保存到JSON文件 with open("hpo_embeddings.json", "w") as f: json.dump(hpo_embeddings, f)
2. 批量请求优化
OpenAI Embedding API支持批量输入,将多个hpo_desc打包成一个请求,减少请求次数,同时提升效率:
import json import openai # 从JSON文件加载HPO字典 with open("hpos.json") as f: hpos_dict = json.load(f) # 创建存储embeddings的字典 hpo_embeddings = {} i = 0 hposNumber = len(hpos_dict) # 批量生成embedding for hpo_id, hpo_descs in hpos_dict.items(): # 直接将所有描述作为批量输入发送请求 response = openai.Embedding.create( input=hpo_descs, model="text-embedding-ada-002" ) # 按顺序提取每个描述对应的embedding embedding_list = [item["embedding"] for item in response["data"]] hpo_embeddings[hpo_id] = embedding_list i = i + 1 print(f"{i}/{hposNumber}") # 将embedding字典保存到JSON文件 with open("hpo_embeddings.json", "w") as f: json.dump(hpo_embeddings, f)
注意:批量输入的总token数不能超过text-embedding-ada-002的8191token限制,若单条描述过长,需拆分或控制批量数量。
3. 添加固定请求间隔
若不想用批量,可在每次请求后添加固定间隔,降低请求频率:
import time # ... 其他代码不变 ... for hpo_desc in hpo_descs: response = openai.Embedding.create( input=hpo_desc, model="text-embedding-ada-002" ) embedding_list.append(response["data"][0]["embedding"]) time.sleep(0.5) # 每次请求后等待0.5秒 # ... 其他代码不变 ...
总结
优先采用批量请求+重试机制的组合方案,既能减少请求次数,又能在服务器过载时自动重试,最大化成功率和效率。
内容的提问来源于stack exchange,提问作者Alberto
相关产品推荐
相关产品推荐

