如何通过存储Toolkit/Agent优化LlamaIndex聊天机器人构建耗时?
解决方案:序列化Agent或预存QueryEngine以加速构建
核心问题分析
你当前的耗时主要来自重复从S3加载130个索引,以及基于索引重复创建QueryEngine、Toolkit和Agent的过程。LlamaIndex的LlamaToolkit本身不支持直接序列化,但可以通过序列化最终的Agent,或预存依赖的QueryEngine来避免重复耗时步骤。
方法一:直接序列化Agent(推荐)
Agent是最终交互的核心,直接序列化Agent可以跳过索引加载、Toolkit构建的全流程,下次直接加载即可使用。
步骤1:修改代码保存Agent
在生成agent_chain后,添加序列化代码:
import pickle def generate_chatbot(): # ... 原有的所有代码 ... llm=ChatOpenAI(temperature=0.1, model="gpt-4") agent_chain = create_llama_chat_agent( toolkit, llm, memory=memory, ) # 序列化Agent到本地磁盘 with open("canada_immigrant_agent.pkl", "wb") as f: pickle.dump(agent_chain, f) return {"agent_chain":agent_chain,"firstMessage":True,"prompt":prompt}
步骤2:快速加载Agent
后续启动时,直接加载序列化的Agent,无需重复加载索引和构建Toolkit:
import pickle from langchain.memory import ConversationBufferMemory from langchain.chat_models import ChatOpenAI def load_chatbot(): # 加载序列化的Agent with open("canada_immigrant_agent.pkl", "rb") as f: agent_chain = pickle.load(f) # 重置会话内存(如果需要清空历史对话) memory = ConversationBufferMemory(memory_key="chat_history") agent_chain.memory = memory prompt = "Hello! You are here to assist me with detailed information for immigrants to Canada. Please refrain from mentioning or discussing any external resources. Instead, kindly provide the relevant content directly and ask for more details if you can't provide relevant content. Don't mention additional sources that are websites. Always ask for more details so that the next answer could be better. Please ensure that the information shared is accurate, up-to-date, and reliable. Don't say to always check the informations on the governement websites. Thank you!" return {"agent_chain":agent_chain,"firstMessage":True,"prompt":prompt}
注意事项
- 序列化文件会包含所有索引和QueryEngine对象,文件体积可能较大,建议存储在本地高速磁盘而非S3。
- 确保LlamaIndex、LangChain、OpenAI等依赖库的版本一致,避免反序列化失败。
- 如果索引内容更新,需要重新生成并序列化Agent。
方法二:预存QueryEngine,快速构建Toolkit
如果不想序列化整个Agent,可以预存每个索引对应的QueryEngine,下次直接加载QueryEngine来构建Toolkit,跳过索引加载后的QueryEngine初始化步骤。
步骤1:保存QueryEngine
在生成custom_query_engines和graph_query_engine后,添加保存代码:
import pickle def generate_chatbot(): # ... 加载索引、构建graph、生成custom_query_engines等代码 ... # 保存每个索引的QueryEngine for idx_id, qe in custom_query_engines.items(): with open(f"query_engine_{idx_id}.pkl", "wb") as f: pickle.dump(qe, f) # 保存Graph的QueryEngine with open("graph_query_engine.pkl", "wb") as f: pickle.dump(graph_query_engine, f) # ... 后续构建Toolkit、Agent的代码 ...
步骤2:加载QueryEngine快速构建Toolkit
后续启动时,直接加载预存的QueryEngine,跳过索引加载后的QueryEngine创建:
import pickle from llama_index import StorageContext, load_index_from_storage from llama_index.tools import IndexToolConfig, LlamaToolkit from llama_index.agent import create_llama_chat_agent from langchain.memory import ConversationBufferMemory from langchain.chat_models import ChatOpenAI def load_chatbot_fast(): urls = [f"canada/{i}" for i in range(130)] index_set = {} custom_query_engines = {} # 加载索引(仍需执行,但可通过本地缓存加速,比如把S3索引同步到本地) for url in urls: # 建议:将S3的索引文件夹同步到本地,替换此处的persist_dir为本地路径 sc_express = StorageContext.from_defaults(persist_dir='myS3/'+url, fs=s3) express_entry_index = load_index_from_storage(sc_express) index_set[url] = express_entry_index # 加载预存的QueryEngine with open(f"query_engine_{express_entry_index.index_id}.pkl", "rb") as f: custom_query_engines[express_entry_index.index_id] = pickle.load(f) # 加载Graph的QueryEngine with open("graph_query_engine.pkl", "rb") as f: graph_query_engine = pickle.load(f) # 快速构建Toolkit graph_config = IndexToolConfig( query_engine=graph_query_engine, name=f"Graph Index", description="useful when asking global questions ", tool_kwargs={"return_direct": True, "return_sources": True}, return_sources=True ) index_configs = [] for url in urls: index = index_set[url] query_engine = custom_query_engines[index.index_id] tool_config = IndexToolConfig( query_engine=query_engine, name=f"Vector Index {url}", description=f"useful for when you want to answer queries about the {url} ", tool_kwargs={"return_direct": True} ) index_configs.append(tool_config) toolkit = LlamaToolkit( index_configs=index_configs+ [graph_config], graph_configs=[graph_config] ) # 构建Agent memory = ConversationBufferMemory(memory_key="chat_history") llm=ChatOpenAI(temperature=0.1, model="gpt-4") agent_chain = create_llama_chat_agent( toolkit, llm, memory=memory, ) prompt = "Hello! You are here to assist me with detailed information for immigrants to Canada. Please refrain from mentioning or discussing any external resources. Instead, kindly provide the relevant content directly and ask for more details if you can't provide relevant content. Don't mention additional sources that are websites. Always ask for more details so that the next answer could be better. Please ensure that the information shared is accurate, up-to-date, and reliable. Don't say to always check the informations on the governement websites. Thank you!" return {"agent_chain":agent_chain,"firstMessage":True,"prompt":prompt}
额外修复:代码中的bug
你的代码中index_configs循环存在一个错误:所有Vector Index的Tool都使用了最后一个索引的QueryEngine,而非对应url的索引。修改后的代码已在方法二中体现,核心是通过index.index_id从custom_query_engines中获取对应QueryEngine。
内容的提问来源于stack exchange,提问作者Hugo Lamoureux
相关产品推荐
相关产品推荐

