You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Ollama Gemma2:2b的简历聊天机器人漏读内容问题排查

简历AI聊天机器人遗漏内容问题排查与改进方案

问题背景

开发一款基于简历的AI聊天机器人,要求完全依据简历内容回复用户查询。测试Gemma2:2b和Mistral模型时,两款模型均出现相同问题:回复会"跳过"简历中的某些部分内容。调整PDF格式后问题有所缓解,但未彻底解决。

可能原因分析

  1. PDF文本提取不完整:PyPDFLoader对复杂格式简历(如分栏布局、表格、特殊字体或扫描件)的文本提取能力有限,导致部分内容未被导入向量数据库。
  2. 文本切分策略不合理:当前采用固定800字符的chunk大小和80字符的重叠,可能将同一类相关内容(如某段工作经历)拆分到多个独立chunk中,检索时无法完整命中。
  3. 向量检索精度不足:仅使用相似度检索且k=5,可能导致部分相关度稍低但关键的chunk未被选中;或embedding函数对简历类文本的适配性不佳。
  4. 提示词约束不足:现有提示词未明确强制模型覆盖所有相关内容,模型可能基于自身倾向选择性忽略部分信息。

具体改进措施

1. 优化PDF文本提取

  • 替换PDF加载器:改用PyMuPDFLoader(对复杂格式支持更好),代码修改如下:
    # 在pdf_handling.py的load_document函数中替换
    from langchain_community.document_loaders import PyMuPDFLoader
    def load_document():
        document_loader = PyMuPDFLoader("Resume.pdf")
        return document_loader.load()
    
  • 若简历是扫描件,需先通过OCR工具(如pytesseract)提取文本,再导入LangChain。
  • 提取后手动核对文本完整性,确认所有段落、表格内容均被正确提取。

2. 调整文本切分策略

  • 采用语义优先的切分规则,针对简历结构设置分隔符:
    def split_documents(documents: list[Document]):
        text_splitter = RecursiveCharacterTextSplitter(
            chunk_size=500,  # 缩小chunk避免拆分完整经历
            chunk_overlap=100,  # 增加重叠保证内容连贯性
            separators=["\n\n", "\n", "。", "!", "?", ". ", "! ", "? ", " ", ""],  # 适配中英文简历结构
            length_function=len,
            is_separator_regex=False,
        )
        return text_splitter.split_documents(documents)
    
  • 或尝试基于简历结构的自定义切分,比如按"工作经历"、"教育背景"等标题拆分文档。

3. 提升向量检索精度

  • 改用MMR检索(最大边际相关性),平衡相似度和内容多样性:
    # 在chatboy.py的query_rag函数中替换检索方式
    results = db.max_marginal_relevance_search(query_text, k=8)  # 增加返回数量
    
  • 更换embedding函数,比如使用更适合长文本的all-mpnet-base-v2:
    # 在get_embedding_function.py中修改
    from langchain_community.embeddings import HuggingFaceEmbeddings
    def get_embedding_function():
        return HuggingFaceEmbeddings(model_name="all-mpnet-base-v2")
    
  • 增加检索返回的chunk数量(如k=8),让模型有更多上下文可选。

4. 强化提示词约束

修改PROMPT_TEMPLATE,明确要求模型必须覆盖所有相关内容:

PROMPT_TEMPLATE = """
你现在需要扮演Maximiliano López Montaño,**必须完全基于提供的简历内容回答问题,不得遗漏任何相关信息,也不得编造未提及的内容**。如果问题涉及多个部分,请逐一说明。

简历内容如下:
{context}

---

请回答以下问题:{question}
"""

5. 其他验证步骤

  • 在pdf_handling.py中添加打印逻辑,确认所有chunk内容正确:
    def add_to_chroma(chunks: list[Document]):
        # ... 原有代码 ...
        for chunk in chunks:
            print(f"Chunk content: {chunk.page_content[:200]}...")  # 打印前200字符验证
    
  • 测试检索结果:输入一个明确指向简历某部分的问题,查看返回的context是否包含该部分内容,定位是检索还是模型输出环节的问题。

相关代码

pdf_handling.py

import argparse
import os
import shutil
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain.schema.document import Document
from get_embedding_function import get_embedding_function
from langchain_chroma import Chroma

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--reset", action = "store_true", help = "Reset the database.")
    args = parser.parse_args()
    if args.reset:
        print("✨ Clearing Database")
        clear_database()

    documents = load_document()
    chunks = split_documents(documents)
    add_to_chroma(chunks)

def load_document():
    document_loader = PyPDFLoader("Resume.pdf")
    return document_loader.load()

def split_documents(documents: list[Document]):
    text_splitter = RecursiveCharacterTextSplitter(
        chunk_size = 800,
        chunk_overlap = 80,
        length_function = len,
        is_separator_regex = False,
    )
    return text_splitter.split_documents(documents)

def add_to_chroma(chunks: list[Document]):
    db_directory = "my_chroma_data"  
    db_path = os.path.join(db_directory, "chroma.sqlite3")  

    os.makedirs(db_directory, exist_ok=True)
    
    db = Chroma(
        persist_directory=db_directory,  
        embedding_function=get_embedding_function()
    )

    chunks_with_ids = calculate_chunk_ids(chunks)

    existing_items = db.get(include=[]) 
    existing_ids = set(existing_items["ids"])
    print(f"Number of existing documents in DB: {len(existing_ids)}")

    new_chunks = []
    for chunk in chunks_with_ids:
        if chunk.metadata["id"] not in existing_ids:
            new_chunks.append(chunk)

    if len(new_chunks):
        print(f"👉 Adding new documents: {len(new_chunks)}")
        new_chunk_ids = [chunk.metadata["id"] for chunk in new_chunks]
        db.add_documents(new_chunks, ids=new_chunk_ids)
    else:
        print("✅ No new documents to add")

def calculate_chunk_ids(chunks):

    last_page_id = None
    current_chunk_index = 0

    for chunk in chunks:
        source = chunk.metadata.get("source")
        page = chunk.metadata.get("page")
        current_page_id = f"{source}:{page}"

        if current_page_id == last_page_id:
            current_chunk_index += 1
        else:
            current_chunk_index = 0

        chunk_id = f"{current_page_id}:{current_chunk_index}"
        last_page_id = current_page_id

        chunk.metadata["id"] = chunk_id

    return chunks

def clear_database():
    db_path = "my_chroma_data/chroma.sqlite3"
    if os.path.exists(db_path):
        shutil.rmtree(os.path.dirname(db_path))  
        print(f"✨ Database at {db_path} has been cleared.")
    else:
        print("⚠️ No database found to clear.")

if __name__ == "__main__":
    main()

chatboy.py

import os
from langchain_chroma import Chroma  # Updated import
from langchain.prompts import ChatPromptTemplate
from langchain_community.llms.ollama import Ollama
from get_embedding_function import get_embedding_function

PROMPT_TEMPLATE = """
You're roleplaying as Maximiliano López Montaño, based on what the resume says.

This is the resume: {context}

---

Answer the question based on the above context: {question}
"""

def main():
    print("Welcome to the chatbot! Type 'exit' to quit.")
    
    while True:
        query_text = input("You: ")
        
        if query_text.lower() == 'exit':
            print("Goodbye!")
            break
        
        response = query_rag(query_text)
        print(f"AI: {response}")

def query_rag(query_text: str):
    embedding_function = get_embedding_function()

    persist_directory = "my_chroma_data"
    db_path = os.path.join(persist_directory, "chroma.sqlite3")

    if not os.path.exists(db_path):
        print("⚠️ Database not found. Please run pdf_handling.py to create the database.")
        return "No data available."

    db = Chroma(persist_directory=persist_directory, embedding_function=embedding_function)

    results = db.similarity_search_with_score(query_text, k=5)

    context_text = "\n\n---\n\n".join([doc.page_content for doc, _score in results])
    prompt_template = ChatPromptTemplate.from_template(PROMPT_TEMPLATE)
    prompt = prompt_template.format(context=context_text, question=query_text)

    model = Ollama(model="gemma2:2b")
    response_text = model.invoke(prompt)

    formatted_response = f"{response_text}"
    return formatted_response

if __name__ == "__main__":
    main()

内容的提问来源于stack exchange,提问作者Maximiliano López

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 10:54:54