You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于LLaMA Index向已有的LanceDB索引添加新文档

问题:如何向LanceDB索引中持续添加PDF文档?

现有可正常运行的RAG系统代码如下:

from llama_index.core import VectorStoreIndex, Settings, StorageContext, Document, SimpleDirectoryReader, \
    load_index_from_storage
from llama_index.vector_stores.lancedb import LanceDBVectorStore
from llama_index.embeddings.huggingface import HuggingFaceEmbedding


vector_store = LanceDBVectorStore(lancedb)
storage_context = StorageContext.from_defaults(vector_store=vector_store)
documents = SimpleDirectoryReader(input_files=["csr.pdf"]).load_data()
index = VectorStoreIndex.from_documents(
        documents, storage_context=storage_context, uri=lancedb
    )
query_engine = index.as_query_engine()
response = query_engine.query("Installation Example for CSR1000V Router")
print(response)

当前问题:后续执行以下代码添加新文档时,会覆盖原有索引数据:

documents = SimpleDirectoryReader(input_files=["new.pdf"]).load_data()
index = VectorStoreIndex.from_documents(
        documents, storage_context=storage_context, uri=lancedb
    )

解决方案

要实现向现有LanceDB索引中追加文档,需加载已有的索引而非重新创建,再调用索引的insert方法添加新文档。具体操作如下:

  1. 初始化与原有索引关联的LanceDB向量存储和存储上下文
  2. 加载已保存的索引
  3. 读取新文档并插入到索引中

代码示例:

# 初始化向量存储(指向已有的LanceDB数据库)
vector_store = LanceDBVectorStore(lancedb)
storage_context = StorageContext.from_defaults(vector_store=vector_store)

# 加载已有的索引
index = load_index_from_storage(storage_context)

# 读取新文档
new_documents = SimpleDirectoryReader(input_files=["new.pdf"]).load_data()

# 将新文档插入到现有索引中
index.insert(new_documents)

# 验证:执行查询测试
query_engine = index.as_query_engine()
response = query_engine.query("查询内容")
print(response)

关键说明

  • load_index_from_storage:该方法用于加载已持久化的索引,避免重新创建时覆盖原有数据。
  • index.insert():此方法会将新文档转换为向量并追加到LanceDB中,完全保留原有索引数据。
  • 确保每次操作时,LanceDBVectorStore初始化使用的是同一个数据库连接/URI,保证操作的是同一套索引数据。

内容的提问来源于stack exchange,提问作者oscar salgado

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 05:33:18