You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GPT-4o令牌超限问题排查:RAG应用429错误及异步加载疑问

RAG应用令牌超限问题排查

问题背景

我运行一个双UI的RAG FAQ聊天机器人:一个UI负责上传PDF文件并嵌入存储到Pinecone向量库,另一个UI从指定索引检索内容用于对话。使用付费Tier1的GPT-4o账号(每分钟30000令牌额度),通过AzureAIDocumentIntelligenceLoader的aload()异步加载272页PDF后,仅输入'hi'就触发令牌超限错误:

'message': 'Request too large for gpt-4o in organization org-wOFxlX2RaRVsbRdbSuZ5iBGM on tokens per min (TPM): Limit 30000, Requested 49634. The input or output tokens must be reduced in order to run successfully.', 'type': 'tokens', 'param': None, 'code': 'rate_limit_exceeded'

改用PyPDFium2Loader加载PDF则对话正常。有两个疑问:

  1. 仅输入'hi'为何会请求近5万令牌?
  2. 已给PDF加载加异步逻辑、检索响应加60秒延迟,为何仍出429错误?

问题解答

疑问1:输入'hi'触发高令牌请求的原因

核心问题出在文档加载与拆分环节:

  • AzureAIDocumentIntelligenceLoader的prebuilt-layout模型会提取PDF中的布局信息(如表格、页眉页脚、格式标记),生成的文本内容比PyPDFium2Loader纯文本加载冗余得多。
  • 你设置的chunk_size=10000过大,单个文本块的令牌数远超预期。加上RAG使用create_stuff_documents_chain,会把检索到的所有文档块直接塞进prompt传给GPT-4o。哪怕输入是'hi',只要检索器返回多个大文本块,prompt总令牌数就会瞬间突破3万限制。
  • 另外检查会话历史:如果StreamlitChatMessageHistory中积累了大量历史对话,也会被计入总令牌数。

疑问2:异步与延迟仍触发429的原因

  • 异步加载不影响令牌消耗:异步只是提升PDF加载效率,和对话时的令牌请求量、速率限制无关。你遇到的是每分钟令牌数(TPM)超限,不是请求频率(RPM)问题。
  • 60秒延迟时机错误:run_query里的time.sleep(60)是在发起请求前等待,但如果单次请求的令牌数直接超过3万,哪怕只发一次,也会触发TPM超限。延迟解决的是请求频率问题,解决不了单请求令牌超标的情况。
  • 可能存在重复请求:检查Streamlit的运行逻辑,是否因为组件重渲染导致多次调用run_query,叠加令牌消耗触发限制。

修复建议

  • 调整文档拆分策略:把chunk_size从10000降到1500左右,chunk_overlap设为200,减少单个文本块的令牌数:
    splt_docs = RecursiveCharacterTextSplitter(chunk_size=1500, chunk_overlap=200)
    
  • 限制检索返回的文档数量:初始化检索器时设置k值,控制每次返回的文档块数量:
    retriever = dbx.as_retriever(search_kwargs={"k": 3})
    
  • 优化文档加载结果:对AzureAIDocumentIntelligenceLoader返回的文本做清洗,去除冗余的布局标记、重复的页眉页脚。
  • 监控会话历史:限制会话历史的长度,比如只保留最近5轮对话,避免历史积累导致令牌超限。
  • 替换链类型:如果文档块还是过大,改用create_map_reduce_chain或create_refine_chain,避免一次性把所有文档塞进prompt。

相关代码

PDF嵌入上传代码

async def extract_embeddings_upload_index(pdf_path, index_name):
    print(f"Loading PDF from path: {pdf_path}")
    
    # Load PDF documents
    async def lol(pdf_path):
        client= await AzureAIDocumentIntelligenceLoader( api_key="167f20e5ce49431aad891c46e2268696",file_path=pdf_path,api_endpoint="https://rx11.cognitiveservices.azure.com/",api_model="prebuilt-layout",mode="single").aload()
        return client

    txt_docs = await lol(pdf_path)
    
    # Split documents
    print("Splitting documents...")
    splt_docs = RecursiveCharacterTextSplitter(chunk_size=10000, chunk_overlap=1000)
    docs = splt_docs.split_documents(txt_docs)
    print(f"Split into {len(docs)} chunks")

    # Initialize OpenAI embeddings
    print("Initializing OpenAI embeddings...")
    embeddings = OpenAIEmbeddings(model='text-embedding-ada-002')

    # Upload documents to Pinecone index
    print("Initializing Pinecone Vector Store...")
    dbx = PineconeVectorStore.from_documents(documents=docs, index_name=index_name, embedding=embeddings)
    print(f"Uploaded {len(docs)} documents to Pinecone index '{index_name}'")

检索链初始化代码

def initialize(index_name):
    embeddings = ini_embed()
    print('11')
    dbx = PineconeVectorStore.from_existing_index(index_name=index_name, embedding=embeddings)
    print('12')
    llm = ChatOpenAI(model='gpt-4o', temperature=0.5, max_tokens=3000)
    
    print('13')
    prompt = ini_prompt()
    print('14')
    doc_chain = create_stuff_documents_chain(llm, prompt)
    print('15')
    retriever = dbx.as_retriever()
    print('16')
    ans_retrieval = create_retrieval_chain(retriever, doc_chain)
    print('17')

    # Wrap the retrieval chain with RunnableWithMessageHistory
    conversational_ans_retrieval = RunnableWithMessageHistory(
        ans_retrieval,
        lambda session_id: StreamlitChatMessageHistory(key=session_id),
        input_messages_key="input",
        history_messages_key="chat_history",
        output_messages_key="answer"
    )
    print('17')
    
    print(session_id)
    print('18')

    return conversational_ans_retrieval

查询执行代码

def run_query(retrieval_chain, input_text):
    st.write('run query')
    try:
        # Generate a response using the retrieval chain
        time.sleep(60)
        response = retrieval_chain.invoke(
            {"input": input_text},
            config={"configurable": {"session_id": f'{session_id}'}}
        )
        
        return response['answer']
    except KeyError as e:
        st.error(f"KeyError occurred: {e}. Check the response structure.")
        return None

内容的提问来源于stack exchange,提问作者Rishil Boddula

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 17:46:04