LangChain中AgentExecutor调用MapReduceDocumentsChain触发上下文超限错误
问题:LangChain工具调用Agent触发摘要工具时出现上下文token超限错误
我在LangChain中整合3个模型,通过OpenAI工具调用Agent根据用户问题调度对应模型。各链已封装为StructuredTool,其中RAG工具和统计工具可正常工作,但调用摘要工具时触发如下错误:
BadRequestError: Error code: 400 - {'error': {'message': "This model's maximum context length is 8192 tokens. However, your messages resulted in 15188 tokens (14729 in the messages, 459 in the functions). Please reduce the length of the messages or functions.", 'type': 'invalid_request_error', 'param': 'messages', 'code': 'context_length_exceeded'}}
单独调用摘要工具无异常,仅通过Agent调用时出现问题。目前摘要工具未使用用户输入,直接传入全部文档,调整整合LLM的token上限后问题仍存在。
摘要工具代码
def summary_maker(): """This function will give back a summary of the emails. You should call this function when the user wants to get a summary of any kind. Also you can call this model if the user question is very broad such as: question: What did my competitors do in the last month? question: What did company X do this last month? """ global docs document_variable_name = "context" llm_summary = load_llm(temp=temp, max_tokens_count=900) # # prompt for first summarizing each document prompt = PromptTemplate.from_template(imp_indiv_email_summarizer_prompt) # # map each document to an individual summary # map_chain = LLMChain(llm=llm_summary, prompt=prompt) # """The ReduceDocumentsChain handles taking the document mapping results " # and reducing them into a single output. It wraps a generic # CombineDocumentsChain (like StuffDocumentsChain) but adds the ability # to collapse documents before passing it to the CombineDocumentsChain if # their cumulative size exceeds token_max. # So if the cumulative number of tokens in our mapped documents exceeds token_max # then we'll recursively pass in the documents in batches of < token_max tokens to our # StuffDocumentsChain to create batched summaries. # see: https://python.langchain.com/v0.2/docs/tutorials/summarization/#map-reduce # """ # # We now define how to combine these summaries reduce_prompt = PromptTemplate.from_template(summary_reduce_prompt) # # and summarize the set of summaries reduce_llm_chain = LLMChain(llm=llm_summary, prompt=reduce_prompt) document_prompt = PromptTemplate( input_variables=["page_content"], template="{page_content}" ) llm_chain = LLMChain(llm=llm_summary, prompt=prompt) combine_documents_chain = StuffDocumentsChain( llm_chain=reduce_llm_chain, document_prompt=document_prompt, document_variable_name=document_variable_name ) # this should be called only if documents exeed context for 'StuffDocumentsChain' collapse_documents_chain = StuffDocumentsChain( llm_chain=llm_chain, document_prompt=document_prompt #document_variable_name=document_variable_name ) # this is the final chain reduce_documents_chain = ReduceDocumentsChain( combine_documents_chain=combine_documents_chain, collapse_documents_chain=collapse_documents_chain, token_max=7000 ) chain = MapReduceDocumentsChain( llm_chain=llm_chain, reduce_documents_chain=reduce_documents_chain, ) # out = chain.invoke({'input_documents' : docs}) # return out # for the record, this is a very bad way of doing this (the fact that we take docs as global parameter and just add it as page_content). # Preferably the page_content is dynamically taken as output of another agent (statistical model) so that you can dynamically get summaries of partial emails def get_summary(input: str) -> str: #print(input) #print("map_red input keys: ", map_reduce_documents_chain.input_keys) print(f"Number of documents: {len(docs)}") result = chain.invoke({"input_documents": docs}) print(f"Summary length: {len(result)}") #print("result:", result) return result summary_tool = StructuredTool.from_function(func=get_summary, description="""This function will give back a summary of the emails. You should call this function when the user wants to get a summary of any kind. Also you can call this model if the user question is very broad such as: question: What did my competitors do in the last month? question: What did company X do this last month? request: Give me a summary of X """) return summary_tool
工具调用Agent初始化代码
prompt3 = PromptTemplate.from_template( """ Answer the following questions as best as you can. You have access to the following tools:\n Semantic email retrieval: Answers a question by converting it into a vector and finds the most similar email. Input should be the question and it will output the answer. Use the following format: Question: The input question you have to answer Thought: you should always think about what to do Action: the action to take should be one of [None, Semantic email retrieval] Observation: the result of the action (this thought/action/observation can repeat N times) Thought: I now know the final answer. Final answer: the final answer to the original input question Begin! Question: {input} Thought: {agent_scratchpad} """ ) # here we add the functions that in the end give us back the answer. tools = [semantic_email_retrieval(), statistical_and_regex_answering(), summary_maker()] llm_integrate = load_llm(0, 1300) # give the tools to a openai_tools_agent # for more info see (https://python.langchain.com/v0.1/docs/modules/agents/agent_types/tool_calling/) agent = create_openai_tools_agent(llm_integrate, tools, prompt3) # the agent executor will be the runtime so to say of this agent. It can for instance route the tasks to the right agent agent_executor = AgentExecutor(agent=agent, tools=tools, return_intermediate_steps=True) output = agent_executor.invoke({"input" : "Please summarize the emails"}) print(output)
解决方案
1. 阻断大内容进入Agent上下文
单独调用摘要工具正常,说明Agent框架会将工具执行的长摘要结果加入对话上下文,导致总token超限。给摘要工具添加return_direct=True参数,让工具结果直接作为Agent最终输出,跳过后续思考环节:
summary_tool = StructuredTool.from_function( func=get_summary, description="""This function will give back a summary of the emails...""", return_direct=True )
2. 限制摘要输出长度
在摘要链的LLM配置中降低max_tokens_count,比如从900调整到500,控制单份摘要的token规模:
llm_summary = load_llm(temp=temp, max_tokens_count=500)
3. 优化Agent Prompt格式
修改Prompt,移除多轮Thought/Action/Observation的循环要求,让Agent调用摘要工具后直接返回结果,避免重复存储大内容:
prompt3 = PromptTemplate.from_template( """ Answer the following questions as best as you can. You have access to the following tools: - Semantic email retrieval: Answers a question by converting it into a vector and finds the most similar email. Input should be the question. - Summary tool: Generates a summary of emails. Use this when asked for summaries or broad questions. If the question requires a summary, call the summary tool directly and return its output as the final answer. For other questions, use the semantic retrieval tool if needed. Begin! Question: {input} Thought: {agent_scratchpad} """ )
4. 调整摘要链的token阈值
降低ReduceDocumentsChain的token_max参数,比如从7000改为5000,给Agent自身的Prompt和工具描述预留足够token空间:
reduce_documents_chain = ReduceDocumentsChain( combine_documents_chain=combine_documents_chain, collapse_documents_chain=collapse_documents_chain, token_max=5000 )
5. 替换大上下文模型
如果以上优化仍不满足需求,将llm_integrate切换到支持更大上下文的模型,比如gpt-3.5-turbo-16k(16k token上限),从根源解决超限问题。
内容的提问来源于stack exchange,提问作者turkishelehant
相关产品推荐
相关产品推荐

