MongoDB Atlas与Langchain预过滤器报错:‘Path needs to be indexed as token’求助
在Langchain与MongoDB Atlas中使用预过滤器的正确配置方案
报错原因解析
你遇到的Path 'question' needs to be indexed as token错误,核心原因是:MongoDB Atlas向量搜索的预过滤逻辑,要求过滤字段必须包含在向量搜索索引的配置中,并且被定义为token类型(而非普通的text索引或单字段索引)。普通的createIndex({question: "text"})无法满足向量搜索预过滤的要求。
解决步骤
1. 更新向量搜索索引,添加question字段为token类型
你需要修改已有的向量搜索索引,将question字段纳入索引配置并设置为token类型,有两种操作方式:
方式一:通过MongoDB Atlas UI操作
- 登录Atlas控制台,进入目标集群,找到对应的数据库(
theme)和集合(ResponseFragment对应的集合,通常是小写复数形式如responsefragments) - 切换到Search标签页,找到你已创建的向量搜索索引(比如
default) - 点击Edit Index,在
mappings.fields数组中添加以下配置:{ "type": "token", "path": "question", "multi": false, "analyzer": "lucene.keyword" } - 保存索引,等待Atlas完成索引重建(时间取决于数据量,通常几分钟)
方式二:通过Python代码创建/更新索引
使用pymongo的create_search_index方法直接配置向量搜索索引:
from pymongo import MongoClient # 替换为你的Atlas连接字符串 client = MongoClient("mongodb+srv://<username>:<password>@<cluster-url>/") db = client["theme"] # 替换为你的ResponseFragment对应的集合名 collection = db["responsefragments"] # 定义向量搜索索引配置,包含embedding、text和question字段 index_config = { "mappings": { "dynamic": False, "fields": { "embedding": { "dimensions": 1536, # 替换为你的embedding实际维度 "similarity": "cosine", "type": "knnVector" }, "text": { "type": "string" }, "question": { "type": "token", "analyzer": "lucene.keyword" } } } } # 创建或更新名为"default"的向量搜索索引 collection.create_search_index(index_config, name="default")
2. 修正代码中的预过滤器参数错误
你的代码中存在变量名不一致的问题:定义了pre_filter_dic但传入的是pre_filter_dict,同时question字段存储的是ObjectId,无需转成字符串,修正后的代码片段如下:
def get_relevant_docs(query, store="theme", k=100, as_dataframe=True, index_name="default", threshold=0.91): # Get the MongoDB collection collection = self.mongo_client[store] # 直接使用ObjectId,无需转字符串 pre_filter_dic = {"question": self.question.id} # Create the MongoDBAtlasVectorSearch instance vectorstore = MongoDBAtlasVectorSearch(collection, self.embeddings, text_key="text", embedding_key="embedding", index_name=index_name) # 传入正确的预过滤器变量 docs = vectorstore.similarity_search_with_score(query, k=k, pre_filter=pre_filter_dic)
3. 验证效果
索引重建完成后,运行修正后的代码,预过滤器会先筛选出关联指定question的ResponseFragment,再在这个子集上执行向量搜索,既满足业务需求又降低性能消耗。
内容的提问来源于stack exchange,提问作者user791793
相关产品推荐
相关产品推荐

