You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用OpenAI API生成Embedding时遭遇服务器过载错误求助

解决OpenAI Embedding请求时的服务器过载错误

问题描述

批量生成HPO描述的Embedding时,收到服务器过载错误:

The server is currently overloaded with other requests. Sorry about that! You can retry your request, or contact us through our help center at help.openai.com if the error persists.

你的代码如下:

# 从JSON文件加载HPO字典
with open("hpos.json") as f:
    hpos_dict = json.load(f)

# 创建存储embeddings的字典
hpo_embeddings = {}

i = 0
hposNumber = len(hpos_dict)
# 生成embedding并存储到字典中
for hpo_id, hpo_descs in hpos_dict.items():
    embedding_list = []
    for hpo_desc in hpo_descs:
        response = openai.Embedding.create(
            input=hpo_desc,
            model="text-embedding-ada-002"
        )
        embedding_list.append(response["data"][0]["embedding"])
    hpo_embeddings[hpo_id] = embedding_list
    i = i + 1
    print( str(i) + "/" + str(hposNumber) )

# 将embedding字典保存到JSON文件
with open("hpo_embeddings.json", "w") as f:
    json.dump(hpo_embeddings, f)

解决方案

这个错误是因为OpenAI服务器负载过高,加上代码采用同步串行请求,短时间内发送大量请求触发了服务器保护机制。可以通过以下几种方式解决:

1. 添加指数退避重试机制

遇到过载错误时自动重试,且每次重试间隔逐渐延长,避免持续给服务器施压。示例代码:

import time
import json
import openai

# 从JSON文件加载HPO字典
with open("hpos.json") as f:
    hpos_dict = json.load(f)

# 创建存储embeddings的字典
hpo_embeddings = {}

i = 0
hposNumber = len(hpos_dict)
# 生成embedding并存储到字典中
for hpo_id, hpo_descs in hpos_dict.items():
    embedding_list = []
    for hpo_desc in hpo_descs:
        retry_count = 0
        max_retries = 5
        while retry_count < max_retries:
            try:
                response = openai.Embedding.create(
                    input=hpo_desc,
                    model="text-embedding-ada-002"
                )
                embedding_list.append(response["data"][0]["embedding"])
                break
            except openai.error.APIError as e:
                if "overloaded" in str(e).lower():
                    retry_count += 1
                    wait_time = 2 ** retry_count  # 指数退避:2,4,8,16,32秒
                    print(f"服务器过载,{wait_time}秒后重试...")
                    time.sleep(wait_time)
                else:
                    raise e
            except Exception as e:
                print(f"未知错误:{e}")
                raise e
    hpo_embeddings[hpo_id] = embedding_list
    i = i + 1
    print(f"{i}/{hposNumber}")

# 将embedding字典保存到JSON文件
with open("hpo_embeddings.json", "w") as f:
    json.dump(hpo_embeddings, f)

2. 批量请求优化

OpenAI Embedding API支持批量输入,将多个hpo_desc打包成一个请求,减少请求次数,同时提升效率:

import json
import openai

# 从JSON文件加载HPO字典
with open("hpos.json") as f:
    hpos_dict = json.load(f)

# 创建存储embeddings的字典
hpo_embeddings = {}

i = 0
hposNumber = len(hpos_dict)
# 批量生成embedding
for hpo_id, hpo_descs in hpos_dict.items():
    # 直接将所有描述作为批量输入发送请求
    response = openai.Embedding.create(
        input=hpo_descs,
        model="text-embedding-ada-002"
    )
    # 按顺序提取每个描述对应的embedding
    embedding_list = [item["embedding"] for item in response["data"]]
    hpo_embeddings[hpo_id] = embedding_list
    i = i + 1
    print(f"{i}/{hposNumber}")

# 将embedding字典保存到JSON文件
with open("hpo_embeddings.json", "w") as f:
    json.dump(hpo_embeddings, f)

注意:批量输入的总token数不能超过text-embedding-ada-002的8191token限制,若单条描述过长,需拆分或控制批量数量。

3. 添加固定请求间隔

若不想用批量,可在每次请求后添加固定间隔,降低请求频率:

import time
# ... 其他代码不变 ...
for hpo_desc in hpo_descs:
    response = openai.Embedding.create(
        input=hpo_desc,
        model="text-embedding-ada-002"
    )
    embedding_list.append(response["data"][0]["embedding"])
    time.sleep(0.5)  # 每次请求后等待0.5秒
# ... 其他代码不变 ...

总结

优先采用批量请求+重试机制的组合方案,既能减少请求次数,又能在服务器过载时自动重试,最大化成功率和效率。

内容的提问来源于stack exchange,提问作者Alberto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 05:22:54