如何在LlamaIndex+ChromaDB中配置BAAI/bge-small-en-v1.5嵌入模型
解决ChromaDB指定BAAI/bge-small-en-v1.5嵌入模型的问题及优化实现
一、手动指定Chroma集合的嵌入函数
直接调用chromadb的get_or_create_collection时,默认会使用内置的all-MiniLM-L6-v2模型,需要显式传入对应模型的嵌入函数,同时要和LlamaIndex Settings里的模型保持一致,避免向量不兼容。
修改后的代码如下:
import chromadb from chromadb.utils.embedding_functions import HuggingFaceEmbeddingFunction from llama_index.vector_stores.chroma import ChromaVectorStore from llama_index.core.storage import StorageContext from llama_index.core.indices import VectorStoreIndex from llama_index.core.settings import Settings from llama_index.embeddings.huggingface import HuggingFaceEmbedding # 配置LlamaIndex全局嵌入模型 Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5") # 创建与LlamaIndex匹配的Chroma嵌入函数 embedding_function = HuggingFaceEmbeddingFunction( model_name="BAAI/bge-small-en-v1.5" ) # 初始化Chroma客户端并创建/获取指定嵌入模型的集合 db = chromadb.PersistentClient(path="<Folder path>") chroma_collection = db.get_or_create_collection( "myCollection", embedding_function=embedding_function ) # 后续构建索引流程不变 vector_store = ChromaVectorStore(chroma_collection=chroma_collection) storage_context = StorageContext.from_defaults(vector_store=vector_store) index = VectorStoreIndex.from_documents(documents=documents, storage_context=storage_context)
注意:必须保证Chroma的embedding_function与LlamaIndex Settings中的模型完全一致,包括模型名称、量化参数等,否则生成的向量维度或特征不匹配,会导致检索失效。
二、更优实现方式(让LlamaIndex自动处理Chroma集合)
无需手动创建chromadb客户端和集合,直接通过LlamaIndex的ChromaVectorStore.from_params方法初始化,它会自动复用Settings中的嵌入模型,代码更简洁且避免手动同步模型的问题。
代码示例:
from llama_index.vector_stores.chroma import ChromaVectorStore from llama_index.core.storage import StorageContext from llama_index.core.indices import VectorStoreIndex from llama_index.core.settings import Settings from llama_index.embeddings.huggingface import HuggingFaceEmbedding # 配置全局嵌入模型 Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5") # 由LlamaIndex自动创建/管理Chroma集合,自动同步嵌入模型 vector_store = ChromaVectorStore.from_params( persist_dir="<Folder path>", collection_name="myCollection" ) storage_context = StorageContext.from_defaults(vector_store=vector_store) index = VectorStoreIndex.from_documents(documents=documents, storage_context=storage_context)
优势
- 自动同步LlamaIndex全局配置的嵌入模型,无需手动维护Chroma的嵌入函数;
- 后续更换嵌入模型时,仅需修改
Settings.embed_model即可,无需调整Chroma相关代码; - 减少冗余代码,降低出错概率。
内容的提问来源于stack exchange,提问作者tigger tigger
相关产品推荐
相关产品推荐

