You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用LlamaIndex调用OpenAI API遭遇RateLimitError问题求助

问题

用户运行以下代码时触发RateLimitError,提示超出当前配额:

import os
import sys

import transformers
from transformers import AutoModelForSequenceClassification, AutoTokenizer

from llama_index import Document, GPTVectorStoreIndex

os.environ['OPENAI_API_KEY'] = 'my-openapi-key'

# Load the hugging face model
model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased")
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# Create a Document object for each text file in the directory
documents = []
for filename in os.listdir("data"):
    with open(os.path.join("data", filename), "r") as f:
        print(filename)
        documents.append(Document(filename, f.read()))

# Create a GPTVectorStoreIndex object from a list of Document objects
index = GPTVectorStoreIndex.from_documents(documents)

# Index the documents
index.index()


# Query the index
query = "What is the capital of France?"
predictions = index.query(query)

# Print the predictions
for prediction in predictions:
    print(prediction)

报错信息:

RateLimitError                            Traceback (most recent call last)
File ~/.local/lib/python3.10/site-packages/tenacity/__init__.py:382, in Retrying.__call__(self, fn, *args, **kwargs)
    381 try:
--> 382     result = fn(*args, **kwargs)
    383 except BaseException:  # noqa: B902

File ~/.local/lib/python3.10/site-packages/llama_index/embeddings/openai.py:149, in get_embeddings(list_of_text, engine, **kwargs)
    147 list_of_text = [text.replace("\n", " ") for text in list_of_text]
--> 149 data = openai.Embedding.create(input=list_of_text, model=engine, **kwargs).data
    150 return [d["embedding"] for d in data]

...(中间报错栈省略)

RetryError: RetryError[<Future at 0x7f6cd45685b0 state=finished raised RateLimitError>]

核心错误提示翻译:超出当前配额,请检查你的套餐和账单详情

解决方法
  • 核查OpenAI配额状态
    查看OpenAI后台的API配额使用情况:若免费额度耗尽,需升级为付费套餐;若付费套餐超出限额,可调整配额额度或等待配额周期重置(通常按小时/月度重置)。

  • 替换为本地Embedding模型(摆脱OpenAI依赖)
    代码中加载了BERT但未实际使用,LlamaIndex默认调用OpenAI Embedding接口生成向量。可改用HuggingFace开源模型本地生成嵌入,无需调用OpenAI API,修改后代码如下:

import os
import sys

from transformers import AutoModelForSequenceClassification, AutoTokenizer
from llama_index import Document, GPTVectorStoreIndex
from llama_index.embeddings import HuggingFaceEmbedding
from llama_index import ServiceContext

# 初始化本地Embedding模型
embed_model = HuggingFaceEmbedding(model_name="bert-base-uncased")
service_context = ServiceContext.from_defaults(embed_model=embed_model)

# Create a Document object for each text file in the directory
documents = []
for filename in os.listdir("data"):
    with open(os.path.join("data", filename), "r") as f:
        print(filename)
        documents.append(Document(filename, f.read()))

# 使用自定义service_context创建索引
index = GPTVectorStoreIndex.from_documents(documents, service_context=service_context)

# Query the index
query = "What is the capital of France?"
predictions = index.query(query)

# Print the predictions
for prediction in predictions:
    print(prediction)
  • 验证API密钥有效性
    确认OPENAI_API_KEY输入正确,密钥拥有访问Embedding接口的权限,未被平台限制使用。

内容的提问来源于stack exchange,提问作者Ankit Bansal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 11:22:37