使用LlamaIndex调用OpenAI API遭遇RateLimitError问题求助
问题
用户运行以下代码时触发RateLimitError,提示超出当前配额:
import os import sys import transformers from transformers import AutoModelForSequenceClassification, AutoTokenizer from llama_index import Document, GPTVectorStoreIndex os.environ['OPENAI_API_KEY'] = 'my-openapi-key' # Load the hugging face model model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased") tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased") # Create a Document object for each text file in the directory documents = [] for filename in os.listdir("data"): with open(os.path.join("data", filename), "r") as f: print(filename) documents.append(Document(filename, f.read())) # Create a GPTVectorStoreIndex object from a list of Document objects index = GPTVectorStoreIndex.from_documents(documents) # Index the documents index.index() # Query the index query = "What is the capital of France?" predictions = index.query(query) # Print the predictions for prediction in predictions: print(prediction)
报错信息:
RateLimitError Traceback (most recent call last) File ~/.local/lib/python3.10/site-packages/tenacity/__init__.py:382, in Retrying.__call__(self, fn, *args, **kwargs) 381 try: --> 382 result = fn(*args, **kwargs) 383 except BaseException: # noqa: B902 File ~/.local/lib/python3.10/site-packages/llama_index/embeddings/openai.py:149, in get_embeddings(list_of_text, engine, **kwargs) 147 list_of_text = [text.replace("\n", " ") for text in list_of_text] --> 149 data = openai.Embedding.create(input=list_of_text, model=engine, **kwargs).data 150 return [d["embedding"] for d in data] ...(中间报错栈省略) RetryError: RetryError[<Future at 0x7f6cd45685b0 state=finished raised RateLimitError>]
核心错误提示翻译:超出当前配额,请检查你的套餐和账单详情
解决方法
核查OpenAI配额状态
查看OpenAI后台的API配额使用情况:若免费额度耗尽,需升级为付费套餐;若付费套餐超出限额,可调整配额额度或等待配额周期重置(通常按小时/月度重置)。替换为本地Embedding模型(摆脱OpenAI依赖)
代码中加载了BERT但未实际使用,LlamaIndex默认调用OpenAI Embedding接口生成向量。可改用HuggingFace开源模型本地生成嵌入,无需调用OpenAI API,修改后代码如下:
import os import sys from transformers import AutoModelForSequenceClassification, AutoTokenizer from llama_index import Document, GPTVectorStoreIndex from llama_index.embeddings import HuggingFaceEmbedding from llama_index import ServiceContext # 初始化本地Embedding模型 embed_model = HuggingFaceEmbedding(model_name="bert-base-uncased") service_context = ServiceContext.from_defaults(embed_model=embed_model) # Create a Document object for each text file in the directory documents = [] for filename in os.listdir("data"): with open(os.path.join("data", filename), "r") as f: print(filename) documents.append(Document(filename, f.read())) # 使用自定义service_context创建索引 index = GPTVectorStoreIndex.from_documents(documents, service_context=service_context) # Query the index query = "What is the capital of France?" predictions = index.query(query) # Print the predictions for prediction in predictions: print(prediction)
- 验证API密钥有效性
确认OPENAI_API_KEY输入正确,密钥拥有访问Embedding接口的权限,未被平台限制使用。
内容的提问来源于stack exchange,提问作者Ankit Bansal
相关产品推荐
相关产品推荐

