GPT-4o令牌超限问题排查:RAG应用429错误及异步加载疑问
RAG应用令牌超限问题排查
问题背景
我运行一个双UI的RAG FAQ聊天机器人:一个UI负责上传PDF文件并嵌入存储到Pinecone向量库,另一个UI从指定索引检索内容用于对话。使用付费Tier1的GPT-4o账号(每分钟30000令牌额度),通过AzureAIDocumentIntelligenceLoader的aload()异步加载272页PDF后,仅输入'hi'就触发令牌超限错误:
'message': 'Request too large for gpt-4o in organization org-wOFxlX2RaRVsbRdbSuZ5iBGM on tokens per min (TPM): Limit 30000, Requested 49634. The input or output tokens must be reduced in order to run successfully.', 'type': 'tokens', 'param': None, 'code': 'rate_limit_exceeded'
改用PyPDFium2Loader加载PDF则对话正常。有两个疑问:
- 仅输入'hi'为何会请求近5万令牌?
- 已给PDF加载加异步逻辑、检索响应加60秒延迟,为何仍出429错误?
问题解答
疑问1:输入'hi'触发高令牌请求的原因
核心问题出在文档加载与拆分环节:
AzureAIDocumentIntelligenceLoader的prebuilt-layout模型会提取PDF中的布局信息(如表格、页眉页脚、格式标记),生成的文本内容比PyPDFium2Loader纯文本加载冗余得多。- 你设置的
chunk_size=10000过大,单个文本块的令牌数远超预期。加上RAG使用create_stuff_documents_chain,会把检索到的所有文档块直接塞进prompt传给GPT-4o。哪怕输入是'hi',只要检索器返回多个大文本块,prompt总令牌数就会瞬间突破3万限制。 - 另外检查会话历史:如果
StreamlitChatMessageHistory中积累了大量历史对话,也会被计入总令牌数。
疑问2:异步与延迟仍触发429的原因
- 异步加载不影响令牌消耗:异步只是提升PDF加载效率,和对话时的令牌请求量、速率限制无关。你遇到的是每分钟令牌数(TPM)超限,不是请求频率(RPM)问题。
- 60秒延迟时机错误:
run_query里的time.sleep(60)是在发起请求前等待,但如果单次请求的令牌数直接超过3万,哪怕只发一次,也会触发TPM超限。延迟解决的是请求频率问题,解决不了单请求令牌超标的情况。 - 可能存在重复请求:检查Streamlit的运行逻辑,是否因为组件重渲染导致多次调用
run_query,叠加令牌消耗触发限制。
修复建议
- 调整文档拆分策略:把
chunk_size从10000降到1500左右,chunk_overlap设为200,减少单个文本块的令牌数:splt_docs = RecursiveCharacterTextSplitter(chunk_size=1500, chunk_overlap=200) - 限制检索返回的文档数量:初始化检索器时设置
k值,控制每次返回的文档块数量:retriever = dbx.as_retriever(search_kwargs={"k": 3}) - 优化文档加载结果:对
AzureAIDocumentIntelligenceLoader返回的文本做清洗,去除冗余的布局标记、重复的页眉页脚。 - 监控会话历史:限制会话历史的长度,比如只保留最近5轮对话,避免历史积累导致令牌超限。
- 替换链类型:如果文档块还是过大,改用
create_map_reduce_chain或create_refine_chain,避免一次性把所有文档塞进prompt。
相关代码
PDF嵌入上传代码
async def extract_embeddings_upload_index(pdf_path, index_name): print(f"Loading PDF from path: {pdf_path}") # Load PDF documents async def lol(pdf_path): client= await AzureAIDocumentIntelligenceLoader( api_key="167f20e5ce49431aad891c46e2268696",file_path=pdf_path,api_endpoint="https://rx11.cognitiveservices.azure.com/",api_model="prebuilt-layout",mode="single").aload() return client txt_docs = await lol(pdf_path) # Split documents print("Splitting documents...") splt_docs = RecursiveCharacterTextSplitter(chunk_size=10000, chunk_overlap=1000) docs = splt_docs.split_documents(txt_docs) print(f"Split into {len(docs)} chunks") # Initialize OpenAI embeddings print("Initializing OpenAI embeddings...") embeddings = OpenAIEmbeddings(model='text-embedding-ada-002') # Upload documents to Pinecone index print("Initializing Pinecone Vector Store...") dbx = PineconeVectorStore.from_documents(documents=docs, index_name=index_name, embedding=embeddings) print(f"Uploaded {len(docs)} documents to Pinecone index '{index_name}'")
检索链初始化代码
def initialize(index_name): embeddings = ini_embed() print('11') dbx = PineconeVectorStore.from_existing_index(index_name=index_name, embedding=embeddings) print('12') llm = ChatOpenAI(model='gpt-4o', temperature=0.5, max_tokens=3000) print('13') prompt = ini_prompt() print('14') doc_chain = create_stuff_documents_chain(llm, prompt) print('15') retriever = dbx.as_retriever() print('16') ans_retrieval = create_retrieval_chain(retriever, doc_chain) print('17') # Wrap the retrieval chain with RunnableWithMessageHistory conversational_ans_retrieval = RunnableWithMessageHistory( ans_retrieval, lambda session_id: StreamlitChatMessageHistory(key=session_id), input_messages_key="input", history_messages_key="chat_history", output_messages_key="answer" ) print('17') print(session_id) print('18') return conversational_ans_retrieval
查询执行代码
def run_query(retrieval_chain, input_text): st.write('run query') try: # Generate a response using the retrieval chain time.sleep(60) response = retrieval_chain.invoke( {"input": input_text}, config={"configurable": {"session_id": f'{session_id}'}} ) return response['answer'] except KeyError as e: st.error(f"KeyError occurred: {e}. Check the response structure.") return None
内容的提问来源于stack exchange,提问作者Rishil Boddula
相关产品推荐
相关产品推荐

