基于Ollama Gemma2:2b的简历聊天机器人漏读内容问题排查
简历AI聊天机器人遗漏内容问题排查与改进方案
问题背景
开发一款基于简历的AI聊天机器人,要求完全依据简历内容回复用户查询。测试Gemma2:2b和Mistral模型时,两款模型均出现相同问题:回复会"跳过"简历中的某些部分内容。调整PDF格式后问题有所缓解,但未彻底解决。
可能原因分析
- PDF文本提取不完整:PyPDFLoader对复杂格式简历(如分栏布局、表格、特殊字体或扫描件)的文本提取能力有限,导致部分内容未被导入向量数据库。
- 文本切分策略不合理:当前采用固定800字符的chunk大小和80字符的重叠,可能将同一类相关内容(如某段工作经历)拆分到多个独立chunk中,检索时无法完整命中。
- 向量检索精度不足:仅使用相似度检索且k=5,可能导致部分相关度稍低但关键的chunk未被选中;或embedding函数对简历类文本的适配性不佳。
- 提示词约束不足:现有提示词未明确强制模型覆盖所有相关内容,模型可能基于自身倾向选择性忽略部分信息。
具体改进措施
1. 优化PDF文本提取
- 替换PDF加载器:改用
PyMuPDFLoader(对复杂格式支持更好),代码修改如下:# 在pdf_handling.py的load_document函数中替换 from langchain_community.document_loaders import PyMuPDFLoader def load_document(): document_loader = PyMuPDFLoader("Resume.pdf") return document_loader.load() - 若简历是扫描件,需先通过OCR工具(如pytesseract)提取文本,再导入LangChain。
- 提取后手动核对文本完整性,确认所有段落、表格内容均被正确提取。
2. 调整文本切分策略
- 采用语义优先的切分规则,针对简历结构设置分隔符:
def split_documents(documents: list[Document]): text_splitter = RecursiveCharacterTextSplitter( chunk_size=500, # 缩小chunk避免拆分完整经历 chunk_overlap=100, # 增加重叠保证内容连贯性 separators=["\n\n", "\n", "。", "!", "?", ". ", "! ", "? ", " ", ""], # 适配中英文简历结构 length_function=len, is_separator_regex=False, ) return text_splitter.split_documents(documents) - 或尝试基于简历结构的自定义切分,比如按"工作经历"、"教育背景"等标题拆分文档。
3. 提升向量检索精度
- 改用MMR检索(最大边际相关性),平衡相似度和内容多样性:
# 在chatboy.py的query_rag函数中替换检索方式 results = db.max_marginal_relevance_search(query_text, k=8) # 增加返回数量 - 更换embedding函数,比如使用更适合长文本的
all-mpnet-base-v2:# 在get_embedding_function.py中修改 from langchain_community.embeddings import HuggingFaceEmbeddings def get_embedding_function(): return HuggingFaceEmbeddings(model_name="all-mpnet-base-v2") - 增加检索返回的chunk数量(如k=8),让模型有更多上下文可选。
4. 强化提示词约束
修改PROMPT_TEMPLATE,明确要求模型必须覆盖所有相关内容:
PROMPT_TEMPLATE = """ 你现在需要扮演Maximiliano López Montaño,**必须完全基于提供的简历内容回答问题,不得遗漏任何相关信息,也不得编造未提及的内容**。如果问题涉及多个部分,请逐一说明。 简历内容如下: {context} --- 请回答以下问题:{question} """
5. 其他验证步骤
- 在pdf_handling.py中添加打印逻辑,确认所有chunk内容正确:
def add_to_chroma(chunks: list[Document]): # ... 原有代码 ... for chunk in chunks: print(f"Chunk content: {chunk.page_content[:200]}...") # 打印前200字符验证 - 测试检索结果:输入一个明确指向简历某部分的问题,查看返回的context是否包含该部分内容,定位是检索还是模型输出环节的问题。
相关代码
pdf_handling.py
import argparse import os import shutil from langchain_community.document_loaders import PyPDFLoader from langchain_text_splitters import RecursiveCharacterTextSplitter from langchain.schema.document import Document from get_embedding_function import get_embedding_function from langchain_chroma import Chroma def main(): parser = argparse.ArgumentParser() parser.add_argument("--reset", action = "store_true", help = "Reset the database.") args = parser.parse_args() if args.reset: print("✨ Clearing Database") clear_database() documents = load_document() chunks = split_documents(documents) add_to_chroma(chunks) def load_document(): document_loader = PyPDFLoader("Resume.pdf") return document_loader.load() def split_documents(documents: list[Document]): text_splitter = RecursiveCharacterTextSplitter( chunk_size = 800, chunk_overlap = 80, length_function = len, is_separator_regex = False, ) return text_splitter.split_documents(documents) def add_to_chroma(chunks: list[Document]): db_directory = "my_chroma_data" db_path = os.path.join(db_directory, "chroma.sqlite3") os.makedirs(db_directory, exist_ok=True) db = Chroma( persist_directory=db_directory, embedding_function=get_embedding_function() ) chunks_with_ids = calculate_chunk_ids(chunks) existing_items = db.get(include=[]) existing_ids = set(existing_items["ids"]) print(f"Number of existing documents in DB: {len(existing_ids)}") new_chunks = [] for chunk in chunks_with_ids: if chunk.metadata["id"] not in existing_ids: new_chunks.append(chunk) if len(new_chunks): print(f"👉 Adding new documents: {len(new_chunks)}") new_chunk_ids = [chunk.metadata["id"] for chunk in new_chunks] db.add_documents(new_chunks, ids=new_chunk_ids) else: print("✅ No new documents to add") def calculate_chunk_ids(chunks): last_page_id = None current_chunk_index = 0 for chunk in chunks: source = chunk.metadata.get("source") page = chunk.metadata.get("page") current_page_id = f"{source}:{page}" if current_page_id == last_page_id: current_chunk_index += 1 else: current_chunk_index = 0 chunk_id = f"{current_page_id}:{current_chunk_index}" last_page_id = current_page_id chunk.metadata["id"] = chunk_id return chunks def clear_database(): db_path = "my_chroma_data/chroma.sqlite3" if os.path.exists(db_path): shutil.rmtree(os.path.dirname(db_path)) print(f"✨ Database at {db_path} has been cleared.") else: print("⚠️ No database found to clear.") if __name__ == "__main__": main()
chatboy.py
import os from langchain_chroma import Chroma # Updated import from langchain.prompts import ChatPromptTemplate from langchain_community.llms.ollama import Ollama from get_embedding_function import get_embedding_function PROMPT_TEMPLATE = """ You're roleplaying as Maximiliano López Montaño, based on what the resume says. This is the resume: {context} --- Answer the question based on the above context: {question} """ def main(): print("Welcome to the chatbot! Type 'exit' to quit.") while True: query_text = input("You: ") if query_text.lower() == 'exit': print("Goodbye!") break response = query_rag(query_text) print(f"AI: {response}") def query_rag(query_text: str): embedding_function = get_embedding_function() persist_directory = "my_chroma_data" db_path = os.path.join(persist_directory, "chroma.sqlite3") if not os.path.exists(db_path): print("⚠️ Database not found. Please run pdf_handling.py to create the database.") return "No data available." db = Chroma(persist_directory=persist_directory, embedding_function=embedding_function) results = db.similarity_search_with_score(query_text, k=5) context_text = "\n\n---\n\n".join([doc.page_content for doc, _score in results]) prompt_template = ChatPromptTemplate.from_template(PROMPT_TEMPLATE) prompt = prompt_template.format(context=context_text, question=query_text) model = Ollama(model="gemma2:2b") response_text = model.invoke(prompt) formatted_response = f"{response_text}" return formatted_response if __name__ == "__main__": main()
内容的提问来源于stack exchange,提问作者Maximiliano López
相关产品推荐
相关产品推荐

