You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Jupyter Notebook中优化llama_index本地模型文档检索速度?

优化LlamaIndex本地文档检索速度的解决方案

一、确认核心组件本地运行状态

  • 确保Ollama服务已本地启动:执行ollama serve命令,且通过ollama pull llama3完成模型本地下载,避免查询时隐式拉取模型。
  • 验证嵌入模型本地缓存:BAAI/bge-base-en-v1.5默认会下载到HuggingFace缓存目录,可手动指定缓存路径确认:
    Settings.embed_model = HuggingFaceEmbedding(
        model_name="BAAI/bge-base-en-v1.5",
        cache_dir="./huggingface_cache"
    )
    

二、避免重复生成嵌入与索引

每次运行代码重新加载文档并生成嵌入是主要耗时点,将索引持久化到本地,后续直接加载即可:

from llama_index.core import StorageContext, load_index_from_storage

# 首次运行:生成并保存索引
index = VectorStoreIndex.from_documents(documents)
index.storage_context.persist(persist_dir="./index_storage")

# 后续运行:直接加载已保存的索引
storage_context = StorageContext.from_defaults(persist_dir="./index_storage")
index = load_index_from_storage(storage_context)

三、优化Ollama LLM推理速度

  • 启用GPU加速:通过ollama info确认Ollama识别到本地GPU,启动时可指定显存分配比例:
    OLLAMA_GPU=100 ollama serve  # 分配100%GPU显存,根据硬件调整
    
  • 调整LLM参数:减少上下文窗口、启用流式输出(感知进度),换用量化模型提升速度:
    # 换用8b量化版本模型,速度更快
    Settings.llm = Ollama(
        model="llama3:8b-instruct-q4_0",
        request_timeout=3600.0,
        temperature=0.0,  # 事实查询设为0,减少无关生成
        context_window=4096,  # 根据需求缩小上下文窗口
        streaming=True  # 流式输出可快速看到响应片段
    )
    

四、优化查询引擎检索逻辑

  • 减少检索片段数量:默认返回Top5相关片段,可减少到Top2~3,降低LLM处理的文本量:
    query_engine = index.as_query_engine(similarity_top_k=2)
    
  • 启用查询缓存:重复查询直接返回结果,避免重复检索与推理:
    from llama_index.core.query_engine import RetrieverQueryEngine
    from llama_index.core.retrievers import VectorIndexRetriever
    from llama_index.core.cache import InMemoryCache
    
    retriever = VectorIndexRetriever(index=index, similarity_top_k=2)
    query_engine = RetrieverQueryEngine.from_args(
        retriever=retriever,
        cache=InMemoryCache()
    )
    

五、加速嵌入模型计算

  • 启用GPU加速嵌入:确保PyTorch支持CUDA,指定设备运行嵌入模型:
    import torch
    Settings.embed_model = HuggingFaceEmbedding(
        model_name="BAAI/bge-base-en-v1.5",
        device="cuda" if torch.cuda.is_available() else "cpu"
    )
    
  • 换用轻量嵌入模型:比如BAAI/bge-small-en-v1.5,速度提升明显,检索精度损失可控。

内容的提问来源于stack exchange,提问作者Kavinila

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 15:27:37