You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

chromadb NoIndexException报错:如何加载已有LangChain向量索引?

解决Chroma索引加载报错chromadb.errors.NoIndexException的问题

问题详情

报错信息:

chromadb.errors.NoIndexException: Index not found, please create an instance before querying

现有LangChain索引文件结构:

tree  langchain/
langchain/
├── chroma-collections.parquet
├── chroma-embeddings.parquet
└── index
    ├── id_to_uuid_cfe8c4e5-8134-4f3d-a120-0510e189004f.pkl
    ├── index_cfe8c4e5-8134-4f3d-a120-0510e189004f.bin
    ├── index_metadata_cfe8c4e5-8134-4f3d-a120-0510e189004f.pkl
    └── uuid_to_id_cfe8c4e5-8134-4f3d-a120-0510e189004f.pkl

1 directory, 6 files

运行以下代码加载索引后,执行index.query('hello')触发上述报错:

if __name__ == "__main__":
    ABS_PATH = os.path.dirname(os.path.abspath(__file__))
    DB_DIR = os.path.join(ABS_PATH, 'langchain/index/index_cfe8c4e5-8134-4f3d-a120-0510e189004f.bin')
    client_settings = chromadb.config.Settings(
        chroma_db_impl="duckdb+parquet",
        persist_directory=DB_DIR,
        anonymized_telemetry=True
    )
    fp = './all_files.txt'
    embeddings = OpenAIEmbeddings()
    get_vectorstore = lambda: Chroma(
            collection_name="langchain",
            embedding_function=embeddings,
            client_settings=client_settings,
            persist_directory=DB_DIR,
        )

    if not os.path.exists(fp):
        root_dir = "."  # Replace with the desired root directory
        gitignore_path = os.path.join(root_dir, ".gitignore")
        ignored_patterns = get_ignored_patterns(gitignore_path)
        files = []
        concatenated_content = ""
        for file_path in find_files(root_dir, "*.py", ignored_patterns):
            files.append(file_path)
            with open(file_path, "r") as file:
                file_content = file.read()
                file_section = f"# <START> {file_path}\n{file_content}\n# <END> {file_path}\n"
                concatenated_content += file_section
        with open('all_files.txt', 'w') as f:
            f.write(concatenated_content)
        loader = TextLoader(fp)
        docs = []
        for loader in [loader]:
            docs.extend(loader.load())
        splitter = _get_default_text_splitter()
        sub_docs = splitter.split_documents(docs)
        vectorstore = get_vectorstore().from_documents(sub_docs, embeddings, persist_directory='./langchain', collection_name='langchain')
        index = VectorStoreIndexWrapper(vectorstore=get_vectorstore())
    else:
        # the defaults
        index = VectorStoreIndexWrapper(vectorstore=get_vectorstore())
    breakpoint()
    print(index)

运行日志:

python embed.py 
Using embedded DuckDB with persistence: data will be stored in: /home/jm/pycharm_projects/test/langchain/index/index_cfe8c4e5-8134-4f3d-a120-0510e189004f.bin
> /home/jm/pycharm_projects/test/embed.py(78)<module>()
-> print(index)
(Pdb) index.query('hello')
*** chromadb.errors.NoIndexException: Index not found, please create an instance before querying
(Pdb) 

解决方案

核心问题是**persist_directory被错误设置为具体的索引文件路径,而非索引所在的根目录**。Chroma需要读取整个目录下的集合文件(.parquet)和索引子目录(index/),而非单个二进制文件。

修正步骤

  1. 调整DB_DIR路径,指向包含索引文件的langchain目录:
DB_DIR = os.path.join(ABS_PATH, 'langchain')
  1. 确保client_settings、get_vectorstore以及创建vectorstore时的persist_directory都使用这个正确路径。

修正后的关键代码

if __name__ == "__main__":
    ABS_PATH = os.path.dirname(os.path.abspath(__file__))
    # 修正:指向langchain根目录而非具体bin文件
    DB_DIR = os.path.join(ABS_PATH, 'langchain')
    client_settings = chromadb.config.Settings(
        chroma_db_impl="duckdb+parquet",
        persist_directory=DB_DIR,
        anonymized_telemetry=True
    )
    fp = './all_files.txt'
    embeddings = OpenAIEmbeddings()
    get_vectorstore = lambda: Chroma(
            collection_name="langchain",
            embedding_function=embeddings,
            client_settings=client_settings,
            persist_directory=DB_DIR,
        )

    if not os.path.exists(fp):
        # 原有文件处理逻辑保持不变
        root_dir = "."
        gitignore_path = os.path.join(root_dir, ".gitignore")
        ignored_patterns = get_ignored_patterns(gitignore_path)
        files = []
        concatenated_content = ""
        for file_path in find_files(root_dir, "*.py", ignored_patterns):
            files.append(file_path)
            with open(file_path, "r") as file:
                file_content = file.read()
                file_section = f"# <START> {file_path}\n{file_content}\n# <END> {file_path}\n"
                concatenated_content += file_section
        with open('all_files.txt', 'w') as f:
            f.write(concatenated_content)
        loader = TextLoader(fp)
        docs = []
        for loader in [loader]:
            docs.extend(loader.load())
        splitter = _get_default_text_splitter()
        sub_docs = splitter.split_documents(docs)
        # 保持persist_directory一致
        vectorstore = get_vectorstore().from_documents(sub_docs, embeddings, persist_directory=DB_DIR, collection_name='langchain')
        index = VectorStoreIndexWrapper(vectorstore=get_vectorstore())
    else:
        index = VectorStoreIndexWrapper(vectorstore=get_vectorstore())
    breakpoint()
    print(index)

调整后重新运行代码,即可正常加载索引并执行查询操作。

内容的提问来源于stack exchange,提问作者jmunsch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 19:27:04