LlamaIndex构建RAG时无法从文档中检索到目标信息的问题求助
LlamaIndex构建RAG时无法从文档中检索到目标信息的问题求助
我最近在尝试用LlamaIndex搭建RAG系统,遇到了一个特别困惑的问题:我用的是Ollama里的gemma聊天模型和nomic-embed-text嵌入模型,这套组合在LangChain里能正常检索到文档里的信息,给出我想要的回答,但换成LlamaIndex实现后,系统却一直说找不到相关内容,实在搞不懂哪里出问题了。
先给大家看看我在LangChain里能正常运行的代码,它能正确返回文档里的注册时间信息:
from langchain_ollama import OllamaEmbeddings from langchain_community.llms import Ollama from langchain.document_loaders import PyPDFLoader from langchain.text_splitter import RecursiveCharacterTextSplitter from langchain.vectorstores import FAISS from langchain.memory import ConversationBufferMemory from langchain.chains import ConversationalRetrievalChain model = Ollama(model = 'gemma', temperature = 0.1) embedding = OllamaEmbeddings(model = 'nomic-embed-text') raw_documents = PyPDFLoader(path+file).load() text_splitter = RecursiveCharacterTextSplitter(chunk_size=1500, chunk_overlap=200) documents = text_splitter.split_documents(raw_documents) db = FAISS.from_documents(documents, embedding) memory = ConversationBufferMemory(memory_key='chat_history', return_messages = True) query_engine = ConversationalRetrievalChain.from_llm(model, retriever=db.as_retriever(), memory=memory, verbose=True) response = query_engine.run({'question':'when in the registration period'}) print(response) # 输出结果:The registration period is between Jan to Feb.
上面这段代码完全没问题,能准确返回我要的答案。但换成LlamaIndex的实现后,就出问题了,代码如下:
from llama_index.llms.ollama import Ollama from llama_index.embeddings.ollama import OllamaEmbedding from llama_index.core import VectorStoreIndex, SimpleDirectoryReader, Settings from llama_index.core.ingestion import IngestionPipeline from llama_index.core.node_parser import TokenTextSplitter Settings.llm = Ollama(model="gemma", request_timeout=360.0) Settings.embed_model = OllamaEmbedding(model_name="nomic-embed-text") documents = SimpleDirectoryReader(path).load_data(file) index = VectorStoreIndex.from_documents(documents) query_engine = index.as_query_engine() response = query_engine.query("when in the registration period") print(response) # 输出结果:The provided text does not contain information regarding..., so I am unable to answer this query from the given context.
明明用的是相同的模型和文档,LlamaIndex却检索不到信息,我本来预期它能给出和LangChain一样的结果,现在实在找不到问题所在,有没有大佬能帮我排查一下?
备注:内容来源于stack exchange,提问作者HappyFish
相关产品推荐
相关产品推荐

