基于FAISS与LangChain的食谱搜索系统如何实现排除式相似搜索?
我基于FAISS向量存储、LangChain和Hugging Face的sentence-transformers/all-MiniLM-L6-v2嵌入模型构建食谱搜索系统,支持按食材、饮食限制、菜系、餐型等参数查找食谱。我用以下Prompt生成带饮食限制定义的自然语言查询和过滤术语:
query_prompt = ChatPromptTemplate.from_template( """Analyze the following search parameters for a recipe search query: Dietary Restrictions: {dietary_restrictions} Ingredients: {ingredients} Dietary Restrictions: {dietary_restrictions} Cuisine: {cuisine} Meal Type: {meal_type} Transform these parameters into a natural language query that can be used to search for recipes. Include the definition of the dietary restrictions. Provide your analysis in the following format: Query: [Transformed natural language query] Filter Terms: [List of key terms for filtering, including ingredients, dietary restrictions, cuisine, and meal type] """ )
生成的查询示例:"Dairy-free American snack recipes, excluding all products derived from milk such as cheese, yogurt, and butter."
随后执行相似搜索:
search_results = vector_store.similarity_search(query, k=1000) # 先获取前1000条结果
但相似搜索无法正确排除指定术语,比如传入以下参数时:
find_recipes( ingredients=[], dietary_restrictions=["dairy free"], cuisine="American", meal_type="snacks" )
返回结果仍包含含乳制品的食谱,比如“### 3. Milk Peda Recipe”“### 12. Yogurt Cheese (Labneh)”。
需要解决的核心问题:如何修改相似搜索以正确排除查询中‘excluding’部分的术语或概念?是否可实现负匹配?有哪些优化搜索以尊重排除项的思路?
解决方案
1. 先召回后硬过滤
利用Prompt生成的Filter Terms提取明确的排除术语,在相似搜索召回结果后做精准过滤,直接剔除包含禁用术语的食谱。
示例代码:
# 解析Prompt输出,提取排除术语(以 dairy-free 为例) exclude_terms = ["cheese", "yogurt", "butter", "milk"] # 执行初始相似搜索 search_results = vector_store.similarity_search(positive_query, k=1000) # 硬过滤排除含禁用术语的文档 filtered_results = [ doc for doc in search_results if not any(term.lower() in doc.page_content.lower() for term in exclude_terms) ]
优点是实现简单、逻辑直观;缺点是若初始召回结果大多不符合要求,会浪费计算资源。
2. 负嵌入加权实现语义级负匹配
通过嵌入模型生成排除术语的向量,对查询嵌入做加权调整,从语义层面降低含排除概念食谱的相似度评分。
示例代码:
from langchain.embeddings import HuggingFaceEmbeddings from sklearn.metrics.pairwise import cosine_similarity import numpy as np embeddings = HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2") # 生成正查询嵌入(聚焦需要的食谱特征) positive_query = "Dairy-free American snack recipes" positive_embedding = embeddings.embed_query(positive_query) # 生成排除术语的嵌入并取平均值 exclude_terms = ["cheese", "yogurt", "butter", "milk"] exclude_embeddings = [embeddings.embed_query(term) for term in exclude_terms] avg_exclude_embedding = np.mean(exclude_embeddings, axis=0) # 调整查询嵌入:正嵌入减去加权后的负嵌入(权重可根据效果调整) weight = 0.5 adjusted_query_embedding = np.array(positive_embedding) - weight * np.array(avg_exclude_embedding) # 手动计算相似度并排序(适配FAISS自定义查询) doc_embeddings = vector_store.index.reconstruct_n(0, vector_store.index.ntotal) docs = list(vector_store.docstore._dict.values()) similarities = cosine_similarity([adjusted_query_embedding], doc_embeddings)[0] sorted_indices = np.argsort(similarities)[::-1][:1000] filtered_results = [docs[i] for i in sorted_indices]
这种方式能捕捉语义层面的排除关系,但需要反复调整权重参数,避免过度惩罚导致遗漏相关结果。
3. 优化Prompt明确正负匹配规则
修改Prompt,让输出清晰区分「需要匹配的正查询」和「需要排除的术语列表」,避免将排除项混在自然语言查询中,简化后续处理逻辑。
优化后的Prompt:
query_prompt = ChatPromptTemplate.from_template( """Analyze the following search parameters for a recipe search query: Dietary Restrictions: {dietary_restrictions} Ingredients: {ingredients} Cuisine: {cuisine} Meal Type: {meal_type} 1. Generate a natural language query focused on the recipes we WANT (include ingredients, dietary restrictions, cuisine, meal type). 2. List all specific terms/ingredients we need to EXCLUDE based on dietary restrictions. Provide your analysis in the following format: Positive Query: [Natural language query for desired recipes] Exclude Terms: [List of exact terms/ingredients to exclude, e.g., cheese, yogurt, butter] """ )
生成的结果结构更清晰,可结合「正查询召回+排除术语硬过滤」或「负嵌入加权」的方式使用。
4. 结构化元数据过滤(若有条件)
如果食谱文档带有结构化元数据(如dietary_tags、ingredients_list字段),可直接利用FAISS的过滤功能,在搜索阶段就剔除不符合要求的结果。
示例代码:
# 假设文档元数据包含dietary_tags字段,标记了食谱的饮食属性 search_results = vector_store.similarity_search( query="Dairy-free American snack recipes", k=1000, filter={"dietary_tags": {"$in": ["dairy-free"]}} )
这种方式效率最高、结果最准确,但依赖食谱数据的结构化程度。
内容的提问来源于stack exchange,提问作者Rasik

